Search NASA⌕ Search

SEARCH · Search NASA

Results for “Probabilistic learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A Probabilistic Reasoner Based on Bayes Risk for Damage Detection in Structural Systems

Structural health monitoring (SHM) systems are used to inform operation of structural systems subject to loads and environments that may affect their integrity. SHM systems rely on continuous monitoring of the structure to determine its health state. These systems are often coupled with a model of the deployed structure to determine the consequences of changes in the system by forecasting the response to future states. These models, which may be thought of as digital twins, need to be updated to reflect the latest state of the structural system. This work makes use of an uncertainty-aware machine learning model that enforces distance preservation of the original input space to determine deviations from the training data input space distributions. This workflow enables domain shift detection to determine whether damage is present in the structure. The uncertainty metrics generated by this network are then used in a Bayes risk framework to design an optimal damage detector given cost and risk considerations. The approach is demonstrated on a computational example with simulated damage.

Najera-Flores, David [ATA Engineering, Inc.]↗

2009 Space Shuttle Probabilistic Risk Assessment Overview

Loss of a Space Shuttle during flight has severe consequences, including loss of a significant national asset; loss of national confidence and pride; and, most importantly, loss of human life. The Shuttle Probabilistic Risk Assessment (SPRA) is used to identify risk contributors and their significance; thus, assisting management in determining how to reduce risk. In 2006, an overview of the SPRA Iteration 2.1 was presented at PSAM 8 [1]. Like all successful PRAs, the SPRA is a living PRA and has undergone revisions since PSAM 8. The latest revision to the SPRA is Iteration 3. 1, and it will not be the last as the Shuttle program progresses and more is learned. This paper discusses the SPRA scope, overall methodology, and results, as well as provides risk insights. The scope, assumptions, uncertainties, and limitations of this assessment provide risk-informed perspective to aid management s decision-making process. In addition, this paper compares the Iteration 3.1 analysis and results to the Iteration 2.1 analysis and results presented at PSAM 8.

Hamlin, Teri L.↗

Emulation With Uncertainty Quantification of Regional Sea‐Level Change Caused by the Antarctic Ice Sheet

Abstract Projecting regional sea‐level change under various climate‐change scenarios typically involves running forward simulations of the Earth's gravitational, rotational and deformational (GRD) response to ice‐mass change, which requires substantial computational cost if applied to probabilistic frameworks requiring thousands to millions of samples. Here we build emulators of regional sea‐level change at 27 coastal locations, due to the GRD effects associated with future Antarctic Ice Sheet mass change over the 21st century. The emulators are evaluated against a numerical sea‐level model applied to an ensemble of ice‐sheet model simulations of the Antarctic Ice Sheet through 2100. We build a physics‐based emulator using a recent sensitivity kernel approach and compare it to machine learning based emulators (neural network and conditional variational autoencoder methods). In order to quantify uncertainty, we derive well‐calibrated prediction intervals for regional sea‐level change via split‐conformal inference and linear regression, and show that Monte Carlo dropout does not yield well‐calibrated uncertainties in this instance. We also demonstrate substantial gains in computational efficiency using both the physics‐based emulator and neural networks in comparison to the numerical model for the complete regional sea‐level solution. Overall, we find the physics‐based emulator modestly outperforms the machine learning emulators for this problem.

58 GEOSCIENCES↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗

Classifying Unidentified X-Ray Sources in the Chandra Source Catalog Using A Multiwavelength Machine-Learning Approach

The rapid increase in serendipitous X-ray source detections requires the development of novel approaches to efficiently explore the nature of X-ray sources. If even a fraction of these sources could be reliably classified, it would enable population studies for various astrophysical source types on a much larger scale than currently possible. Classification of large numbers of sources from multiple classes characterized by multiple properties (features) must be done automatically and supervised machine learning (ML) seems to provide the only feasible approach. We perform classification of Chandra Source Catalog version 2.0 (CSCv2) sources to explore the potential of the ML approach and identify various biases, limitations, and bottlenecks that present themselves in these kinds of studies. We establish the framework and present a flexible and expandable Python pipeline, which can be used and improved by others. We also release the training data set of 2941 X-ray sources with confidently established classes. In addition to providing probabilistic classifications of 66,369 CSCv2 sources (21% of the entire CSCv2 catalog), we perform several narrower-focused case studies (high-mass X-ray binary candidates and X-ray sources within the extent of the H.E.S.S. TeV sources) to demonstrate some possible applications of our ML approach. We also discuss future possible modifications of the presented pipeline, which are expected to lead to substantial improvements in classification confidences.

Hui Yang↗

Learning earthquake ground motions via conditional generative modeling

Predicting high-fidelity ground motions for future earthquakes is crucial for seismic hazard assessment and infrastructure resilience. Conventional empirical simulations suffer from sparse sensor distribution and geographically localized earthquake locations, while physics-based methods are computationally intensive and require accurate representations of Earth structures and earthquake sources. We propose an artificial intelligence (AI) spectrogram generator, Conditional Generative Modeling for Ground Motion (CGM-GM). CGM-GM leverages earthquake magnitudes and geographic coordinates of earthquakes and sensors as inputs, when postprocessed with phase information, capturing spatially continuous Fourier amplitude spectra (FAS) as well as properties such as P and S arrivals, and waveform durations, without explicit physics constraints. This is achieved through a probabilistic autoencoder that extracts latent distributions in the time-frequency domain and variational sequential models for prior and posterior distributions. We evaluate the performance of CGM-GM using small-magnitude earthquake records from the San Francisco Bay Area, a region with high seismic risks. Here, we report that CGM-GM demonstrates potential for complementing physics-based simulations and non-ergodic empirical ground motion models, as well as shows promise in seismology and beyond.

geophysics↗

Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration

This paper proposes novel noise-free Bayesian optimization strategies that rely on a random exploration step to enhance the accuracy of Gaussian process surrogate models. The new algorithms retain the ease of implementation of the classical GP-UCB algorithm, but the additional random exploration step accelerates their convergence, nearly achieving the optimal convergence rate. Furthermore, to facilitate Bayesian inference with intractable likelihoods, we propose to utilize optimization iterates for maximum a posteriori estimation to build a Gaussian process surrogate model for the unnormalized log-posterior density. We provide bounds for the Hellinger distance between the true and the approximate posterior distributions in terms of the number of design points. We demonstrate the effectiveness of our Bayesian optimization algorithms in nonconvex benchmark objective functions, in a machine learning hyperparameter tuning problem, and in a black-box engineering design problem. The effectiveness of our posterior approximation approach is demonstrated in two Bayesian inference problems for parameters of dynamical systems.

Bayesian inference↗

Resolving Mixtures of Soot Characterized by SP-AMS Spectra Using a Latent Dirichlet Allocation Model

Soot produced by detonation or combustion events exhibits different chemical properties depending on the fuel, device construction, and environmental conditions in which the event occurs. These properties can be useful for defining relevant signatures for probabilistically identifying the different types of events that occurred, based on the soot that is produced from these events. However, it is rare to observe samples of soot from a detonation or combustion that are not contaminated by outside particles. In this paper, we present a method for resolving mixtures of soot to determine the contributions of sources that may be present in samples of recovered soot. We use Latent Dirichlet Allocation to describe the generative process for a sample of recovered soot, and use Variational Bayesian Inference to learn about the parameters associated with the generative model. We demonstrate the utility of this method by considering real samples of mixtures of soot under various frameworks to show that the model is able to identify the different components present in a sample of soot as well as their mixing proportions.

54 ENVIRONMENTAL SCIENCES↗

Gateway Program Safety and Mission Assurance Integration - the Future of Safe Deep Space Human Exploration

As a foundational element of the National Aeronautics and Space Administration (NASA) Artemis Campaign, the Gateway is an incrementally built cislunar spacecraft that will serve as a platform for deep space human exploration, science, and technology demonstration. The Gateway will be a unifying catalyst for international partners around the world to establish sustained deep space scientific investigations, lunar surface access, and missions to Mars. As human exploration moves farther away from Earth, spacecraft designs must prioritize and optimize mass and volume allocations, while minimizing human and spacecraft risk. To accomplish this objective, the Gateway Program Safety and Mission Assurance functions develop, implement, and ensure compliance with requirements, in concert with the accurate characterization and transparent communication of residual hazard risks, for integrated safety, reliability and maintainability and quality assurance. Safety and Mission Assurance was a key contributor during Gateway program pre-formulation and formulation activities where safety and reliability analysis was embedded in the Gateway Systems Engineering and Integration team. During these early program stages, a preliminary Gateway Integrated Hazard Analysis and Preliminary Gateway Probabilistic Risk Assessment assisted in Gateway architectural and operational definition as part of a risk-informed design process. As the deep space architecture has matured, the integrated Safety and Mission Assurance analyses have matured, new safety review processes have been developed, and requirements have been refined to ensure compliance with integrated safety and mission assurance objectives. The Gateway Program is currently concluding the preliminary design review informed milestone, where the primary objectives included: - Ensured completeness and consistency of the preliminary design, including the meeting of all requirements within appropriate margins and acceptable risk posture. - Identification of any major issues moving forward to the Critical Design phase. At this milestone, Safety and Mission Assurance provided numerous products, including Gateway Top Risks and Risk Mitigation Plans, updated integrated hazard analyses, updated probabilistic risk assessment, Crew Survival Analysis Report, and updated Safety and Mission Assurance Requirements and Plans. These products provide a many-faceted perspective on the inherent risk and available mitigations involved in flying the current proposed vehicle design and anticipated stack configurations. In addition, Safety and Mission Assurance identified top technical, process and workforce concerns to be addressed as the program progresses toward the critical design phase. This paper will detail the evolution of the Gateway Program Safety and Mission Assurance integration functions, provide its current status and lessons learned for future human spaceflight programs. Throughout this paper the key tenets of the Gateway Program Safety and Mission Assurance will be discussed: - Application of a risk-informed approach to identify and mitigate areas of highest risk. - Leverage of valuable processes and lessons learned from earlier spaceflight programs. - Development of Safety and Mission Assurance products to inform design risk trades. - Utilization of common Safety and Mission Assurance practices to identify safety risks for multiple perspectives: top-down, bottom-up, and across lines of integration. - Approval of safety hazards at the appropriate level of authority, keeping most deliberation closest to design expertise and elevating risks of greatest concern for program-level consideration. - Championing of Safety and Mission Assurance processes and forums to foster a pervasive safety culture that is transparent, inclusive, and collaborative between all partners. These tenets have allowed the Gateway Safety and Mission Assurance function to play a key role in optimized vehicle design evolution, and early identification and mitigation of Gateway program and Artemis mission risk.

Helen Vaccaro↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Efficient Calibration of Expensive Computational Models

Accounting for uncertainty when calibrating expensive computational models is a common challenge faced by scientists and engineers. Often Bayesian techniques are adopted to estimate a probability density function over the model parameters given noisy empirical data. The methods used to perform this type of probabilistic calibration are computationally prohibitive in that they require a large number of evaluations of the expensive model. In these cases, surrogate modeling -- that is, using a fast-to-evaluate, lower fidelity stand-in for the original computational model -- may be the only option to alleviate this computational burden. However, the upfront cost of generating training data to build a surrogate model can itself be expensive. As such, it is important to be judicious when selecting training points at which the full-fidelity model is evaluated. Here, an active learning approach is proposed that enables efficient selection of training points using approximate samples of the calibrated parameter probability density function. In this way, the training points can be concentrated in regions where the calibration algorithm requires high model accuracy.

active learning↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

The composite load spectra project

Probabilistic methods and generic load models capable of simulating the load spectra that are induced in space propulsion system components are being developed. Four engine component types (the transfer ducts, the turbine blades, the liquid oxygen posts and the turbopump oxidizer discharge duct) were selected as representative hardware examples. The composite load spectra that simulate the probabilistic loads for these components are typically used as the input loads for a probabilistic structural analysis. The knowledge-based system approach used for the composite load spectra project provides an ideal environment for incremental development. The intelligent database paradigm employed in developing the expert system provides a smooth coupling between the numerical processing and the symbolic (information) processing. Large volumes of engine load information and engineering data are stored in database format and managed by a database management system. Numerical procedures for probabilistic load simulation and database management functions are controlled by rule modules. Rules were hard-wired as decision trees into rule modules to perform process control tasks. There are modules to retrieve load information and models. There are modules to select loads and models to carry out quick load calculations or make an input file for full duty-cycle time dependent load simulation. The composite load spectra load expert system implemented today is capable of performing intelligent rocket engine load spectra simulation. Further development of the expert system will provide tutorial capability for users to learn from it.

Newell, J. F.↗

Crowd Sourcing Medical Data Collection Using Medical Students

OBJECTIVE We undertook an upgrade of the Evidence Library database of NASA HRP’s Integrated Medical Model, assessing 120 medical conditions which integrate with a novel probabilistic risk assessment (IMPACT) tool of medical risk and resource utilization for long duration exploration human spaceflight. This data collection process included a selection of these conditions crowd sourced over one year via three 4-week medical student electives at the University of Colorado School of Medicine (IDPT 8059 Space Medicine: Human Spaceflight Factors & Medical Risk Assessment). Students undertook a rapid systematic review of each medical condition, under close preceptors with backgrounds in clinical medicine, library science, epidemiology, biostatistics, and evidence-based medicine. As part of the elective, students also received instruction in core space medicine concepts, evidence based medicine and problem based learning sessions as a flight surgeon supporting a simulated Mars mission. METHODS The list of 120 medical conditions includes both common, terrestrial illness/injury (epistaxis, diverticulitis) as well as spaceflight-specific ones (space adaptation conditions, EVA-related injuries). A rapid systematic review process was developed that would allow students to find the data for determining disease incidence/prevalence, return to definitive care (often a surrogate such as hospitalization rates), loss of crew life, and treatment duration. Each data point required a tailored, specialized search process using different databases and corresponding specialized search filters. Databases were selected on their ability to provide high quality literature in an efficient manner and prioritized by their ability to provide graded evidence via a set rubrics specific to human spaceflight. Students were responsible for performing all literature searches and identifying the highest quality available evidence for each data point. Completed student data sheets underwent initial review by faculty preceptors followed by a secondary editing review by the ExMC Clinical Science Team. RESULTS Over the course of three electives, approximately 105 medical conditions were researched by students using spreadsheets with pre-crafted search strategies. Overall, this process was successful in allowing students to perform the preponderance of work to update incidence, treatment duration, return to definitive care, and loss of crew life data points. Students were successful in running searches, identifying the necessary data points within the literature, and determining the types of terrestrial data that most aligns with the astronaut population for successful completion of their tasks. Limitations included variable student experience with search methodologies [PubMed], differing values of evidence grading [best practice evidence based medicine vs. relevant to spaceflight], and students’ unfamiliarity with spaceflight specific conditions. CONCLUSION Finding the relevant literature for medical conditions in spaceflight within terrestrial databases in a systematic method is time consuming and not intuitive. However, the stepwise process that balanced sensitivity with specificity allowed for students to be highly successful in a short amount of time. Additionally, as the process was refined over the course of three electives, preceptors were better able to anticipate where students were likely to encounter barriers, which allowed the course to be adjusted to account for certain data points needing more time for completion. This replicable process may be an efficient way to accomplish rapid systematic reviews for a large volume of data in a short amount of time.

J Lemery↗

VRN3P: Variational Recurrent Neural Network Based Net-Load Prediction under High Solar Penetration

This is the final technical report for the SETO-funded VRN3P project (PNNL# 76914). The goal of this project, led by Pacific Northwest National Laboratory (PNNL), in collaboration with Lawrence Livermore National Laboratory (LLNL) and Portland General Electric (PGE), was to develop and validate a deep variational recurrent neural network-based net-load prediction (VRN3P) framework for probabilistic time-series forecasting of day-ahead net-load under high solar penetration scenarios. The project team reports successful design of a novel probabilistic net-load forecasting architecture, comprising of a variational autoencoder and a recurrent neural network, which demonstrates 30% improvement in forecast performance, 60% improvement in training time, and consumes 44% less memory, when compared with conventional baseline models. The team tested the VRN3P model performance on GridLAB-D test-cases representing varying BTM solar penetration levels of 20%, 30%, and 50%, with integrated time-series net-load profiles provided by the utility partner (PGE). The VRN3P model demonstrate <2% hourly MAPE (averaged over the year) for day- ahead net-load forecast on the test scenario with 20% BTM solar. Transfer learning extension of the VRN3P model has demonstrated 8.33× speed-up in training, while still achieving acceptable forecast performance of 2.24% hourly MAPE on the 30% BTM solar penetration test-scenario. A preliminary version of the VRN3P GridAPPS-D™has been developed, along with a web-based interactive user-interface (named ‘Forte’) which has made available on GitHub for public use.

24 POWER TRANSMISSION AND DISTRIBUTION↗

What Can We Learn From Proton Recoils about Heavy-Ion SEE Sensitivity?

The fact that protons cause single-event effects (SEE) in most devices through production of light-ion recoils has led to attempts to bound heavy-ion SEE susceptibility through use of proton data. Although this may be a viable strategy for some devices and technologies, the data must be analyzed carefully and conservatively to avoid over-optimistic estimates of SEE performance. We examine the constraints that proton test data can impose on heavy-ion SEE susceptibility.

probabilistic risk assessment↗

AUTOMATIC GENERATION OF EVENT TREES AND FAULT TREES: A MODEL-BASED APPROACH

In the past few decades, increasing complexity in modern engineering systems has been driven by the integration of a large number of components and by the fact that the system operations involve many disciplines (e.g., thermal-hydraulics, plant operations, cyber-security). Current safety/reliability modeling approaches to such systems are labor intensive, difficult to learn, and rely heavily on simplistic Boolean logic to depict failure propagation and accident progression. While these methods serve well for simple systems (i.e., linear causal systems with limited small inter- and intra-system interactions), their results are difficult to verify when modeling complex systems (typically performed through the extensive use of modeling assumptions). The development of new methods is addressed to meet these challenges through a model-based system engineering (MBSE) lens. Under MBSE philosophy, every aspect of the system (form or function) is represented by a model that completely characterizes its architecture or behavior. MBSE approach greatly improves the management of design, analysis and verification of complex systems. An integration of Dynamic Probabilistic Risk Assessment (DPRA) methods with MBSE models is proposed to perform safety/reliability analyses of engineering systems. In particular, MBSE representation of the system (performed using Systems Modeling Language [SysML]) is coupled with DPRA methods to automatically generate event trees and fault trees.

97 - MATHEMATICS AND COMPUTING↗

GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism

Graph neural networks (GNNs), an emerging class of machine learning models for graphs, have gained popularity for their superior performance in various graph analytical tasks. Mini-batch training is commonly used to train GNNs on large graphs, and data parallelism is the standard approach to scale mini-batch training across multiple GPUs. Data parallel approaches contain redundant work as subgraphs sampled by different GPUs contain significant overlap. To address this issue, we introduce a hybrid parallel mini-batch training paradigm called Split parallelism. Split parallelism avoids redundant work by splitting the sampling, loading, and training of each mini-batch across multiple GPUs. Split parallelism, however, introduces communication overheads that can be more than the savings from removing redundant work. We further present a lightweight partitioning algorithm that probabilistically minimizes these overheads. We implement spllit parllelism in GSplit and show that it outperforms state-of-the-art mini-batch training systems like DGL, Quiver, and P3.

Lim, Seung-Hwan [ORNL] (ORCID:0000000194616866)↗