Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Classification of Cloud Particle Imagery and Thermodynamics (COCPIT): A New Databasing Tool for the Characterization of Cloud Particle Images Captured During DOE Field Campaigns

The Department of Energy for decades has explored the earth system and atmosphere through research and deployment of in-situ and remote sensing platforms during field campaigns. Among these datasets exists a vast supply of cloud particle images that provide visual insight into the complex microphysics in the clouds that span our globe. The millions of images collected over decades of deployments provides a unique opportunity to further our understanding of our atmosphere down to the crystal size. This work over the past 5 years has sought to organize these images into digestible datasets that can then be used by scientists to further our understanding of microphysics. A machine learning model was developed that categorizes over 1.5 million images across 11 weather events with over 90% accuracy according to particle type. The database was then extended to include dimensional characteristics of the particle as well as co-location of environmental properties, such as temperature and water content. Then, to initialize the connection between these data and our understanding of how crystals form and grow, weather research and forecasting simulations were run to generate the growth histories of the classified crystals. This research culminates with 2 databases per event: (1) a database of all classified crystals and their dimensional and environmental properties and (2) simulated growth histories of each crystal. Finally, a user interface was created to allow researchers to explore data statistics.

54 ENVIRONMENTAL SCIENCES↗

Bayesian prior construction for uncertainty quantification in first-principles statistical mechanics

First-principles statistical mechanics enables the prediction of thermodynamic and kinetic properties of materials, but is computationally expensive. Many approaches require surrogate models to calculate energies within Monte Carlo or molecular dynamics simulations. Inexpensive surrogates such as cluster expansions enable otherwise intractable calculations by interpolating data from higher accuracy methods, such as Density Functional Theory (DFT). Surrogate models introduce uncertainty into downstream calculations, in addition to any uncertainty inherent to DFT calculations. Bayesian frameworks address this by quantifying uncertainty and incorporating expert knowledge through priors. However, constructing effective priors remains challenging. This work introduces and describes practical strategies for building Bayesian cluster expansions, focusing on basis truncation, hyperparameter selection, and ground state replication. We analyze multiple basis truncation schemes, compare cross-validation to the evidence-approximation for hyperparameter optimization, and provide methods to find and enforce ground-state-preserving models through priors. Additionally, we compare the uncertainties between different approximations to DFT (LDA, PBE, SCAN) against the uncertainty introduced with the use of cluster expansion surrogate models. These approaches are demonstrated on the BCC Li x Mg 1-x and Li x Al 1-x alloys, which are both of interest for solid-state Li batteries. Our results provide guidelines for constructing and utilizing Bayesian cluster expansions, thereby improving the transparency of materials modeling. Furthermore, the approaches and insights developed in this work can be transferred to a wide range of cluster expansion surrogate models, including the atomic cluster expansion and related machine-learned interatomic potential architectures.

Alloy theory↗

High-Performance Reaction Wheel Optimization for Fine-Pointing Space Platforms: Minimizing Induced Vibration Effects on Jitter Performance plus Lessons Learned from Hubble Space Telescope for Current and Future Spacecraft Applications

The Hubble Space Telescope (HST) applies large-diameter optics (2.5-m primary mirror) for diffraction-limited resolution spanning an extended wavelength range (approx. 100-2500 nm). Its Pointing Control System (PCS) Reaction Wheel Assemblies (RWAs), in the Support Systems Module (SSM), acquired an unprecedented set of high-sensitivity Induced Vibration (IV) data for 5 flight-certified RWAs: dwelling at set rotation rates. Focused on 4 key ratios, force and moment harmonic values (in 3 local principal directions) are extracted in the RWA operating range (0-3000 RPM). The IV test data, obtained under ambient lab conditions, are investigated in detail, evaluated, compiled, and curve-fitted; variational trends, core causes, and unforeseen anomalies are addressed. In aggregate, these values constitute a statistically-valid basis to quantify ground test-to-test variations and facilitate extrapolations to on-orbit conditions. Accumulated knowledge of bearing-rotor vibrational sources, corresponding harmonic contributions, and salient elements of IV key variability factors are discussed. An evolved methodology is presented for absolute assessments and relative comparisons of macro-level IV signal magnitude due to micro-level construction-assembly geometric details/imperfections stemming from both electrical drive and primary bearing design parameters. Based upon studies of same-size/similar-design momentum wheels' IV changes, upper estimates due to transitions from ground tests to orbital conditions are derived. Recommended HST RWA choices are discussed relative to system optimization/tradeoffs of Line-Of-Sight (LOS) vector-pointing focal-plane error driven by higher IV transmissibilities through low-damped structural dynamics that stimulate optical elements. Unique analytical disturbance results for orbital HST accelerations are described applicable to microgravity efforts. Conclusions, lessons learned, historical context/insights, and perspectives on future applications are given; these previously unpublished data and findings represents a valuable resource for fine-pointing spacecraft or space-based platforms using RWAs, Control Moment Gyros (CMGs), Momentum Wheels, or other ball-bearing-based rotational units.

Hasha, Martin D.↗

Generalizable, fast, and accurate DeepQSPR with fastprop

Abstract Quantitative Structure–Property Relationship studies (QSPR), often referred to interchangeably as QSAR, seek to establish a mapping between molecular structure and an arbitrary target property. Historically this was done on a target-by-target basis with new descriptors being devised to specifically map to a given target. Today software packages exist that calculate thousands of these descriptors, enabling general modeling typically with classical and machine learning methods. Also present today are learned representation methods in which deep learning models generate a target-specific representation during training. The former requires less training data and offers improved speed and interpretability while the latter offers excellent generality, while the intersection of the two remains under-explored. This paper introduces , a software package and general Deep-QSPR framework that combines a cogent set of molecular descriptors with deep learning to achieve state-of-the-art performance on datasets ranging from tens to tens of thousands of molecules. provides both a user-friendly Command Line Interface and highly interoperable set of Python modules for the training and deployment of feedforward neural networks for property prediction. This approach yields improvements in speed and interpretability over existing methods while statistically equaling or exceeding their performance across most of the tested benchmarks. is designed with Research Software Engineering best practices and is free and open source, hosted at github.com/jacksonburns/fastprop.

Burns, Jackson W. (ORCID:0000000206579426)↗

Advanced Bayesian Method for Planetary Surface Navigation

Autonomous Exploration, Inc., has developed an advanced Bayesian statistical inference method that leverages current computing technology to produce a highly accurate surface navigation system. The method combines dense stereo vision and high-speed optical flow to implement visual odometry (VO) to track faster rover movements. The Bayesian VO technique improves performance by using all image information rather than corner features only. The method determines what can be learned from each image pixel and weighs the information accordingly. This capability improves performance in shadowed areas that yield only low-contrast images. The error characteristics of the visual processing are complementary to those of a low-cost inertial measurement unit (IMU), so the combination of the two capabilities provides highly accurate navigation. The method increases NASA mission productivity by enabling faster rover speed and accuracy. On Earth, the technology will permit operation of robots and autonomous vehicles in areas where the Global Positioning System (GPS) is degraded or unavailable.

Center, Julian↗

Application of Support Vector Regression to Derive Crater Depth/Diameter From Satellite Images

Through the study of impact crater shapes, one can draw important conclusions about the nature and evolution of planetary surfaces [e.g., 1-4].In particular, studying the depth (d) to diameter (D)ratio (d/D) of a population of impact craters, in combination with crater count statistics, can yield valuable insights regarding rates of erosion and burial[5]. Motivated by the great abundance of available planetary surface image data, the goal of this project is to develop an efficient way to estimate d/D from satellite images of impact craters for which stereo information is not available [6]. We set out to develop and train a machine learning algorithm to extract d/D from a dataset of synthetic impact crater images for which model d/D is known. The applications of machine learning to planetary science are numerous and diverse [7], including automatic planetary surface mapping [8] and the detection of impact craters [9]. Our algorithm makes use of Support Vector Regression (SVR), which is a type of Support Vector Machine (SVM) [10, 11].SVMs are a branch of supervised machine learning valued for their straightforward implementation and versatility in solving both classification and regression problems. In regression analysis, an SVR algorithm produces a hyperplane function to fit the training data points, as well as an ε-tube that surrounds the hyperplane. Tunable hyperparameters include the width of the ε-tube (ε) and the amount an algorithm is penalized for points which fall outside the ε-tube.

L R Chin↗

The Evaluation of Machine Learning Techniques for Isotope Identification Contextualized by Training and Testing Spectral Similarity

Precise gamma-ray spectral analysis is crucial in high-stakes applications, such as nuclear security. Research efforts toward implementing machine learning (ML) approaches for accurate analysis are limited by the resemblance of the training data to the testing scenarios. The underlying spectral shape of synthetic data may not perfectly reflect measured configurations, and measurement campaigns may be limited by resource constraints. Consequently, ML algorithms for isotope identification must maintain accurate classification performance under domain shifts between the training and testing data. To this end, four different classifiers (Ridge, Random Forest, Extreme Gradient Boosting, and Multilayer Perceptron) were trained on the same dataset and evaluated on twelve other datasets with varying standoff distances, shielding, and background configurations. A tailored statistical approach was introduced to quantify the similarity between the training and testing configurations, which was then related to the predictive performance. Wilcoxon signed-rank tests revealed that the OVR-wrapped XGB significantly outperformed the other algorithms, with confidence levels of 99.0% or above for the 133Ba, 60Co, 137Cs, and 152Eu sources. The findings from this work are significant as they outline techniques to promote the development of robust ML-based approaches for isotope identification.

domain adaptation↗

Generation of random geological models using multi-randomization for machine learning

Generating high-fidelity geological models is essential for advancing machine learning (ML) methods in automated seismic interpretation. For instance, seismic images paired with corresponding fault labels are foundational for ML-based fault detection from seismic migration sections. While several open-access datasets of random geological models exist, open-source tools specifically designed to produce large volumes of such models for ML applications remain scarce. To address this gap, we present RGM (Random Geological Model), an open-source software package for efficiently generating 2D and 3D synthetic geological models tailored for ML workflows. RGM supports the creation of diverse model components, including medium property distributions (P-/S-wave velocities and density), seismic reflectivity images (i.e., synthetic migration sections), relative geological time, and discrete fault attributes such as probability, dip, strike, rake, and displacement. It also accommodates the creation of complex geological features such as salt bodies and unconformities. The model generation algorithm employs a multi-randomization strategy, yielding an effectively infinite-dimensional model space that encompasses a wide range of geological scenarios and associated seismic features. Furthermore, RGM incorporates a method to generate synthetic elastic migration images using analytical elastic reflection coefficients combined with frequency-dependent scaling. This functionality enables the creation of training datasets for ML models that leverage elastic seismic images. RGM is implemented in modern object-oriented Fortran, allowing users to flexibly control statistical parameters governing model variability. We demonstrate the capability, performance, and geological realism of the package through comprehensive 2D and 3D examples.

58 GEOSCIENCES↗

Advanced Method Optimization for Sampling and Analysis Instrumentation

This work presents a generalized approach for analytical method optimization that branches the gap between techniques historically employed and accurate modern optimization techniques suitable for various applications. The novelty of the described strategy is the utilization of multivariate, multiobjective optimization with Karush-Kuhn-Tucker conditions to bound the optimization space to solutions within the physical limitations of instrumentation. Briefly, the basic steps outlined in this paper are to (1) determine the objective(s) that should be maximized or minimized based on the goals of the analytical application, (2) conduct a screening experiment, (3) perform ANOVA to determine the parameters which have a statistically significant effect on the objective, (4) conduct an experiment (e.g., Box-Behnken design) to collect data for fitting the objective equation, and (5) determine the physical constraints of the parameters and solve the Lagrangian to determine the optimal method parameters. A broad approach to optimization target selection allows for robust method tuning to develop improved data sets amenable for chemometrics and machine learning algorithm development. Gas chromatography-mass spectrometry was selected as a use case due to its broad use across scientific fields and time-consuming method development involving numerous parameters. In conclusion, this strategy can reduce the cost of research, improve data quality, and enable the rapid development of new analytical technique.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Minimization of Disorder as a Key Design Principle for Natural Sizes of Light Harvesting 2 Complexes

The light harvesting 2 (LH2) complex of purple bacteria has excellent energy conversion efficiency. Clarifying the design principle behind such efficiency at the atomistic level is crucial for understanding its structure–function relationship and can be utilized for the design of artificial light harvesting systems. To this end, we conducted comprehensive computational investigation of the dynamical and statistical nature of electronic excited states of pigment molecules in a natural LH2 complex with 9-fold symmetry and its two non-natural in silico analogues with 6- and 12-fold symmetries. To ensure reliable and efficient all-atomistic molecular dynamics simulations, we combined a well established interpolation approach for the construction of the potential energy surface with a neural network machine learning approach. Outcomes of these calculations clarify that non-natural forms of LH2-type complexes have significantly larger quasistatic disorder than those for the natural one. In addition, non-natural systems have more disruptions of the hydrogen bonding, underscoring its crucial role for reducing the disorder. On the other hand, local environmental dynamics are relatively insensitive to the structural changes although there is moderate enhancement in the anharmonic or interatomic components for the synthetic ones. These findings based on all-atomistic simulations provide direct computational evidence that the structure and sizes of natural LH2 complexes are designed to minimize the energetic disorder. We analyze quantitative implications of these for the energy transferring capability of the LH2 complex.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Global SO 2 Data Record from OMPS Instruments on the JPSS Constellation

NASA’s Earth Observing System (EOS) SO 2 climate data record (CDR) started in 2004, with the launch of the Aura/Ozone Monitoring Instrument (OMI) and is now being continued with the SNPP/Ozone Mapping and Profiler Suite (OMPS) launched in 2011. Both OMI and SNPP/OMPS SO 2 CDRs are produced with the Goddard principal component analysis (PCA) spectral fitting algorithm. An advantage of the data-driven PCA retrieval technique is that it enables highly consistent retrievals from different instruments, by inherently accounting for various instrumental factors. To further extend the EOS SO 2 CDR, we are implementing the PCA SO 2 retrieval algorithm with the L1B measurements from OMPS instruments flying on the Joint Polar Satellite System (JPSS) constellation. In this presentation, we will provide an update on our progress in NOAA-20 (launched in 2017) and NOAA-21 (launched in 2022) PCA SO2 retrievals. We will focus on our new NOAA-20/OMPS PCA SO 2 EOS continuity product, to be publicly released in fall of 2023. We will present statistical analyses on the quality of NOAA-20 PCA SO 2 product, including retrieval noise, biases over background areas, and long-term stability. We will compare our PCA SO 2 retrievals from NOAA-20 with those from OMI, SNPP/OMPS, and S5P/TROPOMI (TROPOspheric Monitoring Instrument) for anthropogenic sources as well as large volcanic plumes. We will also discuss the application of a new machine learning technique that helps to further reduce the noise of NOAA-20 SO 2 retrievals. In addition, we will present preliminary PCA SO 2 retrievals from NOAA-21/OMPS, including those from direct readout implementation for aviation disaster avoidance. Finally, we will share some first results applying the PCA algorithm to NASA’s geostationary TEMPO (Tropospheric Emissions: Monitoring of Pollution) instrument to obtain hourly, high resolution SO 2 data over North America.

SO2↗

Uncertainty quantification for inverse problems with application to ptychographic reconstruction

Inverse problems in imaging are commonly solved by optimization or learned surrogates that return a single reconstruction, while uncertainty information is often unavailable. In many experimental settings, however, uncertainty is required to assess reliability, guide downstream analysis, and prioritize additional measurements. In this note, we present a compact uncertainty-quantification framework based on local objective curvature, and then specialize it to ptychographic reconstruction. We further show how repeated reconstructions can be aggregated in a statistically principled way, including a practical implementation path for PtychoNN.

97 MATHEMATICS AND COMPUTING↗

Identifying Topological Defects in Lamellar Phases through Contour Analysis of Complex Wave Fields

Lamellar phases frequently contain structural imperfections that significantly affect their behaviors and properties. Our previous research successfully reconstructed real-space configurations of defective lamellar phases from diffuse scattering patterns, indicating the presence of phase vortices as a potential method for identifying topological defects disrupting the smectic ordering. Here, this report presents a mathematical framework using regularized wave fields to represent defective lamellar structures in real space. Phase singularities, resulting from the interference of random waves and indicating lamellar order disruption, are identified through a contour integral. These wave fields, derived from coherent scattering in reciprocal space, were validated via computational benchmarks analyzing small-angle neutron scattering data from AOT surfactant solutions, facilitating further statistical analysis of the defects. Our study highlights the potential to extract meaningful information about topological defects in lyotropic phases by inversely analyzing experimentally measured two-point static correlations. Our method allows for detailed structural analysis of various lyotropic phases, both particulate and nonparticulate, in their quiescent states and facilitates quantitative investigation of defects’ role in phase transitions. By integrating small-angle scattering, deep learning, and vortex tangle analysis, our comprehensive approach shows promise in addressing complex challenges in the structural analysis of soft matter systems.

36 MATERIALS SCIENCE↗

NiCd cell reliability in the mission environment

This paper summarizes an effort by Gates Aerospace Batteries (GAB) and the Reliability Analysis Center (RAC) to analyze survivability data for both General Electric and GAB NiCd cells utilized in various spacecraft. For simplicity sake, all mission environments are described as either low Earth orbital (LEO) or geosynchronous Earth orbit (GEO). 'Extreme value statistical methods' are applied to this database because of the longevity of the numerous missions while encountering relatively few failures. Every attempt was made to include all known instances of cell-induced-failures of the battery and to exclude battery-induced-failures of the cell. While this distinction may be somewhat limited due to availability of in-flight data, we have accepted the learned opinion of the specific customer contacts to ensure integrity of the common databases. This paper advances the preliminary analysis reported upon at the 1991 NASA Battery Workshop. That prior analysis was concerned with an estimated 278 million cell-hours of operation encompassing 183 satellites. The paper also cited 'no reported failures to date.' This analysis reports on 428 million cell hours of operation emcompassing 212 satellites. This analysis also reports on seven 'cell-induced-failures.'

Denson, William K.↗

Exploring and Analyzing Climate Variations Online by Using NASA MERRA-2 Data at GES DISC

NASA Giovanni (Goddard Interactive Online Visualization ANd aNalysis Infrastructure) (http:giovanni.sci.gsfc.nasa.govgiovanni) is a web-based data visualization and analysis system developed by the Goddard Earth Sciences Data and Information Services Center (GES DISC). Current data analysis functions include Lat-Lon map, time series, scatter plot, correlation map, difference, cross-section, vertical profile, and animation etc. The system enables basic statistical analysis and comparisons of multiple variables. This web-based tool facilitates data discovery, exploration and analysis of large amount of global and regional remote sensing and model data sets from a number of NASA data centers. Long term global assimilated atmospheric, land, and ocean data have been integrated into the system that enables quick exploration and analysis of climate data without downloading, preprocessing, and learning data. Example data include climate reanalysis data from NASA Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2) which provides data beginning in 1980 to present; land data from NASA Global Land Data Assimilation System (GLDAS), which assimilates data from 1948 to 2012; as well as ocean biological data from NASA Ocean Biogeochemical Model (NOBM), which provides data from 1998 to 2012. This presentation, using surface air temperature, precipitation, ozone, and aerosol, etc. from MERRA-2, demonstrates climate variation analysis with Giovanni at selected regions.

knowledge base↗

Effective Defect Detection Using Instance Segmentation for NDI

Ultrasonic testing is a common Non-Destructive Inspection (NDI) method used in aerospace manufacturing. However, the complexity and size of the ultrasonic scans make it challenging to identify defects through visual inspection or machine learning models. Using computer vision techniques to identify defects from ultrasonic scans is an evolving research area. In this study, we used instance segmentation to identify the presence of defects in the ultrasonic scan images of composite panels that are representative of real components manufactured in aerospace. We used two models based on Mask- RCNN (Detectron 2) and YOLO 11 respectively. Additionally, we implemented a simple statistical pre-processing technique that reduces the burden of requiring custom-tailored pre-processing techniques. Our study demonstrates the feasibility and effectiveness of using instance segmentation in the NDI pipeline by significantly reducing data pre-processing time, inspection time, and overall costs.

computer vision techniques↗

Mark 4A project training evaluation

A participant evaluation of a Deep Space Network (DSN) is described. The Mark IVA project is an implementation to upgrade the tracking and data acquisition systems of the dSN. Approximately six hundred DSN operations and engineering maintenance personnel were surveyed. The survey obtained a convenience sample including trained people within the population in order to learn what training had taken place and to what effect. The survey questionnaire used modifications of standard rating scales to evaluate over one hundred items in four training dimensions. The scope of the evaluation included Mark IVA vendor training, a systems familiarization training seminar, engineering training classes, a on-the-job training. Measures of central tendency were made from participant rating responses. Chi square tests of statistical significance were performed on the data. The evaluation results indicated that the effects of different Mark INA training methods could be measured according to certain ratings of technical training effectiveness, and that the Mark IVA technical training has exhibited positive effects on the abilities of DSN personnel to operate and maintain new Mark IVA equipment systems.

Stephenson, S. N.↗

Proceedings of the Workshop on Change of Representation and Problem Reformulation

The proceedings of the third Workshop on Change of representation and Problem Reformulation is presented. In contrast to the first two workshops, this workshop was focused on analytic or knowledge-based approaches, as opposed to statistical or empirical approaches called 'constructive induction'. The organizing committee believes that there is a potential for combining analytic and inductive approaches at a future date. However, it became apparent at the previous two workshops that the communities pursuing these different approaches are currently interested in largely non-overlapping issues. The constructive induction community has been holding its own workshops, principally in conjunction with the machine learning conference. While this workshop is more focused on analytic approaches, the organizing committee has made an effort to include more application domains. We have greatly expanded from the origins in the machine learning community. Participants in this workshop come from the full spectrum of AI application domains including planning, qualitative physics, software engineering, knowledge representation, and machine learning.

Lowry, Michael R.↗