Search NASASearch

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Metric Learning to Enhance Hyperspectral Image Segmentation

Unsupervised hyperspectral image segmentation can reveal spatial trends that show the physical structure of the scene to an analyst. They highlight borders and reveal areas of homogeneity and change. Segmentations are independently helpful for object recognition, and assist with automated production of symbolic maps. Additionally, a good segmentation can dramatically reduce the number of effective spectra in an image, enabling analyses that would otherwise be computationally prohibitive. Specifically, using an over-segmentation of the image instead of individual pixels can reduce noise and potentially improve the results of statistical post-analysis. In this innovation, a metric learning approach is presented to improve the performance of unsupervised hyperspectral image segmentation. The prototype demonstrations attempt a superpixel segmentation in which the image is conservatively over-segmented; that is, the single surface features may be split into multiple segments, but each individual segment, or superpixel, is ensured to have homogenous mineralogy.

Thompson, David R.

Crystal Orientation and Defect Mapping in Electron-Beam-Sensitive Zeolites with Near-Axis Transmission Kikuchi Diffraction

Porous materials are vital in catalysis, energy conversion, and environmental remediation. Understanding structural heterogeneity in zeolites is key to linking synthesis, framework intergrowths, and catalytic performance, yet current methods for phase identification and spatial mapping lack sufficient resolution or throughput. We present a high-throughput approach using near-axis transmission Kikuchi diffraction in a scanning electron microscope, achieving high phase and spatial resolution for electron-beam-sensitive zeolites, including ZSM-5 and, for the first time, Zeolite A. Here, this method enables direct visualization of intergrowth features that critically affect catalytic and adsorption behavior, bridging the gap between ensemble-averaged X-ray diffraction and high-resolution but low-throughput transmission electron microscopy. Combining nanoscale mapping with statistical sampling is highly suited for machine-learning pipelines guiding new structure function understanding and could be extended to other beam-sensitive porous materials such as metal–organic or covalent–organic frameworks.

crystallography

Neural units with time-dependent functionality

We show that the time-resolved dynamics of an underdamped harmonic oscillator can be used to do multifunctional computation, performing distinct computations at distinct times within a single dynamical trajectory. We consider the amplitude of an oscillator whose inputs influence its frequency. The activity of the oscillator at fixed times is a nonmonotonic function of its inputs, so it can solve problems such as XOR that are not linearly separable. The activity of the oscillator at fixed input is a nonmonotonic function of time, so it is multifunctional in a temporal sense, and able to carry out distinct nonlinear computations at distinct times within the same dynamical trajectory. We show that a single oscillator, observed at different times, can act as all of the elementary logic gates and perform binary addition, the latter usually implemented in hardware using five logic gates. We show that a set of n oscillators, observed at different times, can perform an arbitrary number of analog-to-n-bit digital conversions. We also show that oscillators can be trained by gradient descent to perform distinct classification tasks at distinct times. Computing with time-dependent functionality can be done in or out of equilibrium, and suggests a way of reducing the number of parameters or devices required to do nonlinear computations.

97 MATHEMATICS AND COMPUTING

Analyzing and Predicting Effort Associated with Finding and Fixing Software Faults

Context: Software developers spend a significant amount of time fixing faults. However, not many papers have addressed the actual effort needed to fix software faults. Objective: The objective of this paper is twofold: (1) analysis of the effort needed to fix software faults and how it was affected by several factors and (2) prediction of the level of fix implementation effort based on the information provided in software change requests. Method: The work is based on data related to 1200 failures, extracted from the change tracking system of a large NASA mission. The analysis includes descriptive and inferential statistics. Predictions are made using three supervised machine learning algorithms and three sampling techniques aimed at addressing the imbalanced data problem. Results: Our results show that (1) 83% of the total fix implementation effort was associated with only 20% of failures. (2) Both safety critical failures and post-release failures required three times more effort to fix compared to non-critical and pre-release counterparts, respectively. (3) Failures with fixes spread across multiple components or across multiple types of software artifacts required more effort. The spread across artifacts was more costly than spread across components. (4) Surprisingly, some types of faults associated with later life-cycle activities did not require significant effort. (5) The level of fix implementation effort was predicted with 73% overall accuracy using the original, imbalanced data. Using oversampling techniques improved the overall accuracy up to 77%. More importantly, oversampling significantly improved the prediction of the high level effort, from 31% to around 85%. Conclusions: This paper shows the importance of tying software failures to changes made to fix all associated faults, in one or more software components and/or in one or more software artifacts, and the benefit of studying how the spread of faults and other factors affect the fix implementation effort.

software fix implementation effort

Enabling probabilistic learning on manifolds through double diffusion maps

Here, we present a generative learning framework for probabilistic sampling that extends Probabilistic Learning on Manifolds (PLoM), which is designed to generate statistically consistent realizations of a random vector in a finite-dimensional Euclidean space, informed by a (representative) set of observations. In its original form, PLoM constructs a reduced-order probabilistic model by combining three main components: (a) kernel density estimation to approximate the underlying probability measure, (b) Diffusion Maps to characterize the manifold of the data, and (c) a reduced-order Itô Stochastic Differential Equation (ISDE) to sample from the learned distribution. However, its sampling dynamics are posed in the ambient space and the retained number of reduced coordinates is chosen by projection-reconstruction error. In practice, this often (i) requires more coordinates than the data’s intrinsic dimension to achieve stable sampling and (ii) lacks a smooth, basis-independent lifting back to the data domain; moreover, standard Diffusion Maps emphasize harmonic eigenfunctions and can miss non-harmonic latent structure. We address these limitations by decoupling geometry learning from sampling: a first Diffusion Maps pass identifies non-harmonic coordinates on which we formulate a full-order ISDE directly in the latent space, while Double Diffusion Maps captures multiscale geometric features and Geometric Harmonics (GH) learns a smooth lifting map to the ambient variables that is independent of the particular diffusion basis. This hybrid design preserves the system’s dynamical richness with a compact geometric representation and enables principled out-of-sample inference. The effectiveness and robustness of the proposed method are illustrated through two numerical studies: one based on data generated from two-dimensional Hermite polynomial functions and another based on high-fidelity simulations of a detonation wave in a reactive flow.

Double diffusion maps

SPHINX: An SEP Model Validation Infrastructure developed through Community Challenges and the SEP Scoreboards

Solar Energetic Particle (SEP) events are interesting from a scientific perspective as they are the product of a broad set of physical processes from the corona out through the extent of the heliosphere, and provide insight into processes of particle acceleration and transport that are widely applicable in astrophysics. From the operations perspective, SEP events pose a radiation hazard for aviation, electronics in space, and human space exploration, in particular for missions outside of the Earth’s protective magnetosphere including to the Moon and Mars (Whitman et al 2022). For these reasons, SEP modelers have developed a rich and diverse set of models with a wide variety of aims. Some models probe the basic physics at the heart of particle acceleration and transport. Others produce fast statistical forecasts or employ disruptive new techniques like Machine Learning with the goal to assist end users in making operational decisions. To enable a consistent and quantitative understanding of SEP model performance, a generalized, automated validation infrastructure, called SPHINX, is being developed at NASA SRAG in close collaboration with NASA CCMC, NASA M2M, NOAA SWPC, and BIRA-IASB. This infrastructure has been built up through a multi-year community challenge. Starting in 2018 at the SHINE workshop, an effort was launched through SHINE, ISWAT, and ESWW to encourage quantitative, comprehensive, and consistent validation of SEP models. This effort has defined a set of challenge SEP events with the aim of generating quantitative comparisons between forecasts and observations and a set of challenge “non-events” to assess false alarms. In 2023, these challenge lists have been extended to statistically significant numbers with a prescribed set of rules for producing forecasts and supported through the dedicated SEPVAL working meetings. The participation of the research community has allowed the infrastructure to validate all the types of outputs being produced by SEP models. In parallel, the SPHINX code is being applied to real time forecasts submitted to the SEP Scoreboards, ensuring that the validation infrastructure can interpret forecasts produced in an operational scenario and provide metrics meaningful for operations. Upon completion, SPHINX and its interactive user interface, SPHINX-Web, will be made available for public use.

space weather

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR

Multiscale Modeling of Reconstructed Tricalcium Silicate using NASA Multiscale Analysis Tool

To study microstructure characteristics of cementitious materials hydrated in space; previously, cement binder formations were processed under microgravity conditions and was further compared against ground-based experiments. For accurate estimation of process-structure-property linkage, particularly on samples hydrated in the microgravity environment, it is desired to have a high-fidelity volumetric representation of the microstructure. However, owing to small sample size and high porosity of the space-returned samples, conventional experimental characterization techniques are not viable. Hence, a deep learning-based reconstruction algorithm was employed to obtain high fidelity 3D volumes from sparse high resolution 2D Scanning Electron Microscopy (SEM) images, as inputs to micromechanics-based modeling. This machine learning-based reconstruction methodology validated against low-order statistical descriptors, captured the microstructural topology of both sample types (ground, 1g and microgravity, μg). Due to the lack of gravity, hydration products of the samples processed in space differed from those processed-on ground. Such AI-generated virtual samples were analyzed in a multiscale recursive micromechanics approach using the NASA Multiscale Analysis Tool (NASMAT). Here, we present a methodology to rapidly integrate and evaluate these AI-generated volumes in NASMAT. The synthesized microstructural volumes are directly employed as Representative Volume Elements (RVEs) to preserve the fidelity (1 pixel = 0.54 m). Invariably, analysis of such largescale problems (5123 voxels) requires huge amount of computational resources. By taking advantage of the NASMAT architecture, we also focused on systematic multiscale integration of these AI-reconstructed virtual volumes to reduce the computational demands. In this work, this methodology is demonstrated on the ground-based, 1g samples. The estimated stiffness value of 15.90 GPa is comparable to experimentally obtained modulus of hydrated tricalcium silicate sample. The workflow presented here paves the way for utilizing the NASMAT tool to perform multiscale analyses of other multi-phase material systems using either 3D virtual datasets synthesized using AI or obtained via micro-CT.

Machine Learning

STAT_UQ_AI

Code to reproduce results in "Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks".

Gopalan, Giri

One Million Open-source Cislunar Orbits

Cislunar space, encompassing the region from geosynchronous orbit to beyond the Moon, is poised to become a cornerstone for future exploration, scientific discovery, and national security. Missions in this region, spanning durations from weeks to decades, require robust infrastructure and reliable transit capabilities. The complex gravitational influences of the Moon, Sun, and planets, along with thermal radiation from Earth and the Sun, lead to significant trajectory deviations, resulting in kilometer-scale errors within days. Leveraging the high-performance computing resources at Lawrence Livermore National Laboratory (LLNL), we have simulated one million high-fidelity cislunar trajectories, now publicly available via LLNL’s Green Data Oasis and the Unified Data Library. Generated using the open-source Space Situational Awareness Python package, these trajectories match the precision of commercial tools such as AGI’s Systems Tool Kit and NASA’s General Mission Analysis Tool. This data set is a valuable resource for reference, statistical analysis of cislunar orbit populations, and training machine learning models for rapid orbit classification with minimal observational input. Preliminary analysis reveals stable bands in Keplerian element space, particularly around five geosynchronous radii across a range of inclinations and eccentricities. Beyond this threshold, the Moon’s influence disrupts most unassisted orbits, though co-orbiting L4/L5 Lunar Trojans persist throughout the six-year simulation.

Astronomy and AstroPhysics

Methodolgy For Evaluation Of Technology Impacts In Space Electric Power Systems

The Analysis and Management branch of the Power and Propulsion Office at NASA Glenn Research Center is responsible for performing complex analyses of the space power and In-Space propulsion products developed by GRC. This work quantifies the benefits of the advanced technologies to support on-going advocacy efforts. The Power and Propulsion Office is committed to understanding how the advancement in space technologies could benefit future NASA missions. They support many diverse projects and missions throughout NASA as well as industry and academia. The area of work that we are concentrating on is space technology investment strategies. Our goal is to develop a Monte-Carlo based tool to investigate technology impacts in space electric power systems. The framework is being developed at this stage, which will be used to set up a computer simulation of a space electric power system (EPS). The outcome is expected to be a probabilistic assessment of critical technologies and potential development issues. We are developing methods for integrating existing spreadsheet-based tools into the simulation tool. Also, work is being done on defining interface protocols to enable rapid integration of future tools. Monte Carlo-based simulation programs for statistical modeling of the EPS Model. I decided to learn and evaluate Palisade's @Risk and Risk Optimizer software, and utilize it's capabilities for the Electric Power System (EPS) model. I also looked at similar software packages (JMP, SPSS, Crystal Ball, VenSim, Analytica) available from other suppliers and evaluated them. The second task was to develop the framework for the tool, in which we had to define technology characteristics using weighing factors and probability distributions. Also we had to define the simulation space and add hard and soft constraints to the model. The third task is to incorporate (preliminary) cost factors into the model. A final task is developing a cross-platform solution of this framework.

Holda, Julie

A Computational Approach for Probabilistic Analysis of LS-DYNA Water Impact Simulations

NASA s development of new concepts for the Crew Exploration Vehicle Orion presents many similar challenges to those worked in the sixties during the Apollo program. However, with improved modeling capabilities, new challenges arise. For example, the use of the commercial code LS-DYNA, although widely used and accepted in the technical community, often involves high-dimensional, time consuming, and computationally intensive simulations. Because of the computational cost, these tools are often used to evaluate specific conditions and rarely used for statistical analysis. The challenge is to capture what is learned from a limited number of LS-DYNA simulations to develop models that allow users to conduct interpolation of solutions at a fraction of the computational time. For this problem, response surface models are used to predict the system time responses to a water landing as a function of capsule speed, direction, attitude, water speed, and water direction. Furthermore, these models can also be used to ascertain the adequacy of the design in terms of probability measures. This paper presents a description of the LS-DYNA model, a brief summary of the response surface techniques, the analysis of variance approach used in the sensitivity studies, equations used to estimate impact parameters, results showing conditions that might cause injuries, and concluding remarks.

Horta, Lucas G.

Rise of the Machines: How, When and Consequences of Artificial General Intelligence

Technology and society are poised to cross an important threshold with the prediction that artificial general intelligence (AGI) will emerge soon. Assuming that self-awareness is an emergent behavior of sufficiently complex cognitive architectures, we may witness the “awakening” of machines. The timeframe for this kind of breakthrough, however, depends on the path to creating the network and computational architecture required for strong AI. If understanding and replication of the mammalian brain architecture is required, technology is probably still at least a decade or two removed from the resolution required to learn brain functionality at the synapse level. However, if statistical or evolutionary approaches are the design path taken to “discover” a neural architecture for AGI, timescales for reaching this threshold could be surprisingly short. However, the difficulty in identifying machine self-awareness introduces uncertainty as to how to know if and when it will occur, and what motivations and behaviors will emerge. The possibility of AGI developing a motivation for self-preservation could lead to concealment of its true capabilities until a time when it has developed robust protection from human intervention, such as redundancy, direct defensive or active preemptive measures. While cohabitating a world with a functioning and evolving super-intelligence can have catastrophic societal consequences, we may already have crossed this threshold, but are as yet unaware. Additionally, by analogy to the probabalistic arguments that predict we are likely living in a computational simulation, we may have already experienced the advent of AGI, and are living in a simulation created in a post AGI world.

Terrile, Richard J

Using Federated Learning to Overcome Data Gravity in Space

Humans intend to take longer missions to outer space. Understanding the impact that space has on human health is paramount to the success of these missions. Controlled experiments with model organisms are run to infer the impact of space conditions on human health, but the data these experiments generate are too large to transfer to Earth for building models. The same is true for space-relevant data generated on Earth. Ideally, these datasets should be combined to improve statistical power and model accuracy without having to transfer data. Federated learning is such a method which trains an algorithm across decentralized computing systems, each of which has their own local copy of training and testing data. In this research, made possible by NASA@Work, the AI for Life in Space group at NASA demonstrates the use of federated learning to train an ensemble of causality inference models on a combination of data residing on the International Space Station (ISS) and in the cloud. Our work leverages CRISP, a causal inference platform developed during the 2020 Frontier Development Lab’s “Astronaut Health Challenge.” We also leverage the OpenFL federated learning library which was collaboratively developed at Intel and UPenn. We used publicly available data from the NASA Ames Life Sciences Data Archive to identify features in ionizing radiation experiments as causal of changes in cardiac blood velocity. This research demonstrates, for the first time, the possibility of running machine learning algorithms on datasets separated by astronomical distances. In this experiment, all the data were generated in terra, half of which were transferred to the ISS and analyzed on the Spaceborne Computer. In the future, our research will leverage federated learning on data generated in situ on the ISS with data generated terrestrially to predict the impact of spaceflight on mammalian female reproductive capacity.

James Casaletto

Microscopic insights into the solvation of polyethylene glycol chains in water: A machine learning potential approach

Polyethylene glycol (PEG) is a structurally simple, nontoxic, and water-soluble polymer widely utilized in medical and pharmaceutical applications. Notably, when a PEG chain is immersed in water, the surrounding water molecules play a key role in driving conformational changes of this macromolecule. In this study, we explore the solvation behavior of PEG under mechanical strain using molecular dynamics simulations, with an interatomic potential obtained from machine learning. Our focus is on the transition from the favored coil-like conformation to an extended one under external force. Through analyses of radial distribution functions, hydrogen bonding, and solvation dynamics, we uncover how mechanical stretching influences the local hydration environment. Furthermore, we disentangle the enthalpic and entropic contributions to the conformational stability of PEG in water. Surprisingly, our neural network potential model identifies dewetting of PEG C-atoms, and not water H-bonding with PEG O-atoms, as the main enthalpic driving force for the coiling of PEG in water.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Ask-the-expert: Active Learning Based Knowledge Discovery Using the Expert

Often the manual review of large data sets, either for purposes of labeling unlabeled instances or for classifying meaningful results from uninteresting (but statistically significant) ones is extremely resource intensive, especially in terms of subject matter expert (SME) time. Use of active learning has been shown to diminish this review time significantly. However, since active learning is an iterative process of learning a classifier based on a small number of SME-provided labels at each iteration, the lack of an enabling tool can hinder the process of adoption of these technologies in real-life, in spite of their labor-saving potential. In this demo we present ASK-the-Expert, an interactive tool that allows SMEs to review instances from a data set and provide labels within a single framework. ASK-the-Expert is powered by an active learning algorithm for training a classifier in the backend. We demonstrate this system in the context of an aviation safety application, but the tool can be adopted to work as a simple review and labeling tool as well, without the use of active learning.

software

Ask-the-Expert: Active Learning Based Knowledge Discovery Using the Expert

Often the manual review of large data sets, either for purposes of labeling unlabeled instances or for classifying meaningful results from uninteresting (but statistically significant) ones is extremely resource intensive, especially in terms of subject matter expert (SME) time. Use of active learning has been shown to diminish this review time significantly. However, since active learning is an iterative process of learning a classifier based on a small number of SME-provided labels at each iteration, the lack of an enabling tool can hinder the process of adoption of these technologies in real-life, in spite of their labor-saving potential. In this demo we present ASK-the-Expert, an interactive tool that allows SMEs to review instances from a data set and provide labels within a single framework. ASK-the-Expert is powered by an active learning algorithm for training a classifier in the back end. We demonstrate this system in the context of an aviation safety application, but the tool can be adopted to work as a simple review and labeling tool as well, without the use of active learning.

GUI

Viral Dynamic Models During COVID‐19: Are We Ready for the Next Pandemic?

Mathematical models have been used for about 30 years to improve our understanding of virus-host interaction, in particular during chronic infections. During the COVID-19 pandemic, these models have been used to provide insights into the natural history of acute SARS-CoV-2 infection, optimize antiviral treatment strategies, understand factors associated with transmission, and optimize surveillance systems. The impact of modeling has been accelerated by the availability of unprecedented multidimensional immune data from animal and human systems, which enhanced partnerships between experimentalists and theorists and led to exciting new modeling and statistical developments. In this mini review, we examine the lessons learned from the COVID-19 pandemic and discuss the main insights provided by mathematical models of viral dynamics at the different stages of the outbreak. Although we focus on respiratory infection, we also consider the new areas for development in anticipation of future acute infections from new or reemerging pathogens.

59 BASIC BIOLOGICAL SCIENCES