Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

“Understanding Robustness Lottery”: A Geometric Visual Comparative Analysis of Neural Network Pruning Approaches

Deep learning approaches have provided state-of-the-art performance in many applications by relying on large and overparameterized neural networks. However, such networks are very brittle and are difficult to deploy on resource-limited platforms. Model pruning, i.e., reducing the size of the network, is a widely adopted strategy that can lead to a more robust and compact model. Many heuristics exist for model pruning, but our understanding of the pruning process remains limited due to the black-box nature of a neural network model. Empirical studies show that some heuristics improve performance whereas others can make models more brittle. Here, this work aims to shed light on how different pruning methods alter the network’s internal feature representation and the corresponding impact on model performance. To facilitate a comprehensive comparison and characterization of the high-dimensional model feature space, we introduce a visual geometric analysis of feature representations. We evaluated a set of critical geometric concepts decomposed from the commonly adopted classification loss and used them to design a visualization system to compare and highlight the impact of pruning on model performance and feature representation. The proposed tool provides an environment for an in-depth comparison of pruning methods and a comprehensive understanding of how the model responds to common data corruption. By leveraging the proposed visualization, machine learning researchers can reveal the similarities between pruning methods and redundancy in robustness evaluation benchmarks, obtain geometric insights about the differences between pruned models that achieve superior robustness performance, and identify samples that are robust or fragile to model pruning and common data corruption.

Li, Zhimin [Univ. of Utah, Salt Lake City, UT (Uni↗

Scattering-based structural inversion of soft materials via Kolmogorov–Arnold networks

Small-angle scattering techniques are indispensable tools for probing the structure of soft materials. However, traditional analytical models often face limitations in structural inversion for complex systems, primarily due to the absence of closed-form expressions of scattering functions. To address these challenges, we present a machine learning framework based on the Kolmogorov–Arnold Network (KAN) for directly extracting real-space structural information from scattering spectra in reciprocal space. This model-independent, data-driven approach provides a versatile solution for analyzing intricate configurations in soft matter. By applying the KAN to lyotropic lamellar phases and colloidal suspensions—two representative soft matter systems—we demonstrate its ability to accurately and efficiently resolve structural collectivity and complexity. Here, our findings highlight the transformative potential of machine learning in enhancing the quantitative analysis of soft materials, paving the way for robust structural inversion across diverse systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning Neutrino-Nucleus Cross Sections

Neutrino-nucleus scattering cross sections are critical theoretical inputs for long-baseline neutrino oscillation experiments. However, robust modeling of these cross sections remains challenging. For a simple but physically motivated toy model of the DUNE experiment, we demonstrate that an accurate neural-network model of the cross section -- leveraging Standard Model symmetries -- can be learned from near-detector data. We then perform a neutrino oscillation analysis with simulated far-detector events, finding that the modeled cross section achieves results consistent with what could be obtained if the true cross section were known exactly. This proof-of-principle study highlights the potential of future neutrino near-detector datasets and data-driven cross-section models.

Wagman, Michael L. [Fermilab] (ORCID:0000000176701↗

A Semi-Supervised Learning Method for the Identification of Bad Exposures in Large Imaging Surveys

As the data volume of astronomical imaging surveys rapidly increases, traditional methods for image anomaly detection, such as visual inspection by human experts, are becoming impractical. We introduce a machine-learning-based approach to detect poor-quality exposures in large imaging surveys, with a focus on the DECam Legacy Survey (DECaLS) in regions of low extinction (i.e., E ( B − V ) < 0.04 ). Our semi-supervised pipeline integrates a vision transformer (ViT), trained via self-supervised learning (SSL), with a k-Nearest Neighbor (kNN) classifier. We train and validate our pipeline using a small set of labeled exposures observed by surveys with the Dark Energy Camera (DECam). A clustering-space analysis of where our pipeline places images labeled in good and bad categories suggests that our approach can efficiently and accurately determine the quality of exposures. Applied to new imaging being reduced for DECaLS Data Release 11, our pipeline identifies 780 problematic exposures, which we subsequently verify through visual inspection. Being highly efficient and adaptable, our method offers a scalable solution for quality control in other large imaging surveys.

Luo, Yufeng (ORCID:0000000246230683)↗

Revealing EDL-driven reduction mechanisms in binary, ternary, and quaternary fluorinated electrolytes via an integrated MD–DFT–ML framework

Accurately predicting solid electrolyte interphase (SEI) formation requires explicitly resolving the electric double layer (EDL) structure, which deviates significantly from that of the bulk electrolyte. Although an established molecular dynamics (MD) and Density Functional Theory (DFT) framework can model SEI formation by evaluating reduction reactions of local clusters in the EDL, it suffers from a combinatorial computational bottleneck. To overcome this limitation, we introduce a machine-learning-accelerated simulation workflow (MD–DFT–ML), integrating a gradient-boosted regression model trained on EDL composition data to efficiently predict reduction potentials. We apply this framework to seven fluorinated electrolytes comprising fluorinated anions, a fluorinated ester solvent, two types of diluent (ion-solvating ester vs. non-solvating ether), and an FEC additive. The analysis shows that the EDL selectively accumulates cation-binding species; consequently, the non–cation-binding ether diluent rarely enters the EDL and makes minimal contributions to SEI formation. DFT calculations on statistically representative EDL clusters provide reduction potentials and fluorine-release pathways, while the ML model, which substantially reduces the DFT workload, predicts cluster reduction energies with a mean absolute error of 0.1 eV. The combined MD–DFT–ML approach also quantifies contributions from different sources to LiF formation in the SEI. This methodology establishes a generalizable route for multiscale modeling electrolyte and interphase design for next-generation electrochemical energy-storage systems.

DFT-MD-ML workflow↗

Online Electron Reconstruction at CLAS12

Online reconstruction plays a crucial role in monitoring and in real-time analysis of high energy and nuclear physics experiments. A vital aspect of reconstruction algorithms is particle identification, which combines information from various detector components to determine the type of particle. Electron identification is particularly significant in electro-production nuclear physics experiments like the CLAS12 spectrometer at Jefferson Laboratory as it is essential in data recording. A machine learning approach has been developed for CLAS12 experiments to reconstruct and identify electrons by combining raw signals from multiple detector components at the data acquisition level. This method achieves high electron identification purity while maintaining nearly 100% efficiency. Furthermore, the machine learning tools operate at rates exceeding data acquisition speed, enabling the real-time electron reconstruction. This advancement significantly improves online analyses and monitoring capabilities for CLAS12 experiments.

Tyson,, Richard [Thomas Jefferson National Acceler↗

Overview of NASA supported Stirling thermodynamic loss research

NASA is funding research to characterize Stirling machine thermodynamic losses. NASA's primary goal is to improve Stirling design codes to support engine development for space and terrestrial power. However, much of the fundamental data is applicable to Stirling cooling and heat pump applications. The research results are reviewed. Much was learned about oscillating flow hydrodynamics, including laminar/turbulent transition, and tabulated data was documented for further analysis. Now, with a better understanding of the oscillating flow field, it is time to begin measuring the effects of oscillating flow and oscillating pressure level on heat transfer in heat exchanger flow passages and in cylinders.

Tew, Roy C.↗

Overview of NASA supported Stirling thermodynamic loss research

NASA is funding research to characterize Stirling machine thermodynamic losses. NASA's primary goal is to improve Stirling design codes to support engine development for space and terrestrial power. However, much of the fundamental data is applicable to Stirling cooling and heat pump applications. The research results are reviewed. Much was learned about oscillating flow hydrodynamics, including laminar/turbulent transition, and tabulated data was documented for further analysis. Now, with a better understanding of the oscillating flow field, it is time to begin measuring the effects of oscillating flow and oscillating pressure level on heat transfer in heat exchanger flow passages and in cylinders.

Tew, Roy C.↗

Classifying thermodynamic cloud phase using machine learning models

Vertically resolved thermodynamic cloud-phase classifications are essential for studies of atmospheric cloud and precipitation processes. The Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Thermodynamic Cloud Phase (THERMOCLDPHASE) value-added product (VAP) uses a multi-sensor approach to classify the thermodynamic cloud phase by combining lidar backscatter and depolarization, radar reflectivity, Doppler velocity, spectral width, microwave-radiometer-derived liquid water path, and radiosonde temperature measurements. The measured pixels are classified as ice, snow, mixed phase, liquid (cloud water), drizzle, rain, and liq_driz (liquid+drizzle). We use this product as the ground truth to train three machine learning (ML) models to predict the thermodynamic cloud phase from multi-sensor remote sensing measurements taken at the ARM North Slope of Alaska (NSA) observatory: a random forest (RF), a multi-layer perceptron (MLP), and a convolutional neural network (CNN) with a U-Net architecture. Evaluations against the outputs of the THERMOCLDPHASE VAP with 1 year of data show that the CNN outperforms the other two models, achieving the highest test accuracy, F1 score, and mean intersection over union (IOU). Analysis of ML confidence scores shows that ice, rain, and snow have higher confidence scores, followed by liquid, while mixed, drizzle, and liq_driz have lower scores. Feature importance analysis reveals that the mean Doppler velocity and vertically resolved temperature are the most influential data streams for ML thermodynamic cloud-phase predictions. Lidar measurements exhibit lower feature importance due to rapid signal attenuation caused by the frequent presence of persistent low-level clouds at the NSA site. The ML models' generalization capacity is further evaluated by applying them at another Arctic ARM site in Norway using data taken during the ARM Cold-Air Outbreaks in the Marine Boundary Layer Experiment (COMBLE) field campaign. The models demonstrated similar performance to that observed at the NSA site. Finally, we evaluate the ML models' response to simulated instrument outages and signal degradation and show that a CNN U-Net model trained with input channel dropouts performs better when input fields are missing.

ARM Aerial Facility↗

Machine learning neutrino-nucleus cross sections

Neutrino-nucleus scattering cross sections are critical theoretical inputs for long-baseline neutrino oscillation experiments. However, robust modeling of these cross sections remains challenging. For a simple but physically motivated toy model of the DUNE experiment, we demonstrate that an accurate neural-network model of the cross section—leveraging only Standard-Model symmetries—can be learned from near-detector data. We perform a neutrino oscillation analysis with simulated far-detector events, finding that oscillation analysis results enabled by our data-driven cross-section model approach the theoretical limit achievable with perfect prior knowledge of the cross section. We further quantify the effects of flux shape and detector resolution uncertainties as well as systematics from cross-section mismodeling. This proof-of-principle study highlights the potential of future neutrino near-detector datasets and data-driven cross-section models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine Learning Neutrino-Nucleus Cross Sections

Neutrino-nucleus scattering cross sections are critical theoretical inputs for long-baseline neutrino oscillation experiments. However, robust modeling of these cross sections remains challenging. For a simple but physically motivated toy model of the DUNE experiment, we demonstrate that an accurate neural-network model of the cross section—leveraging only Standard-Model symmetries— can be learned from near-detector data. We perform a neutrino oscillation analysis with simulated far-detector events, finding that oscillation analysis results enabled by our data-driven cross-section model approach the theoretical limit achievable with perfect prior knowledge of the cross section. We further quantify the effects of flux shape and detector resolution uncertainties as well as systematics from cross-section mismodeling. This proof-of-principle study highlights the potential of future neutrino near-detector datasets and data-driven cross-section models.

Tame-Narvaez, Karla [Fermilab] (ORCID:000000022249↗

Accelerating multilevel Markov Chain Monte Carlo using machine learning models

Here, this work presents an efficient approach for accelerating multilevel Markov Chain Monte Carlo (MCMC) sampling for large-scale problems using low-fidelity machine learning models. While conventional techniques for large-scale Bayesian inference often substitute computationally expensive high-fidelity models with machine learning models, thereby introducing approximation errors, our approach offers a computationally efficient alternative by augmenting high-fidelity models with low-fidelity ones within a hierarchical framework. The multilevel approach utilizes the low-fidelity machine learning model (MLM) for inexpensive evaluation of proposed samples thereby improving the acceptance of samples by the high-fidelity model. The hierarchy in our multilevel algorithm is derived from geometric multigrid hierarchy. We utilize an MLM to accelerate the coarse level sampling. Training machine learning model for the coarsest level significantly reduces the computational cost associated with generating training data and training the model. We present an MCMC algorithm to accelerate the coarsest level sampling using MLM and account for the approximation error introduced. We provide theoretical proofs of detailed balance and demonstrate that our multilevel approach constitutes a consistent MCMC algorithm. Additionally, we derive the expression for cost reduction due to machine learning model to facilitate cost analysis of the hierarchical sampling algorithm. Our technique is demonstrated on a standard benchmark inference problem in groundwater flow, where we estimate the probability density of a quantity of interest using a four-level MCMC algorithm. Our proposed algorithm accelerates multilevel sampling by a factor of two while achieving similar accuracy compared to sampling using the standard multilevel algorithm.

97 MATHEMATICS AND COMPUTING↗

Machine Learning Models for Mapping Groundwater Pollution Risk: Advancing Water Security and Sustainable Development Goals in Georgia, USA

The widespread use of pesticides, such as atrazine and malathion, in agricultural systems raises significant concerns regarding the contamination of groundwater, which serves as a critical resource for drinking water. This study applies machine learning techniques to predict the concentrations of atrazine and malathion in groundwater across Georgia, USA, using 2019 data. A Random Forest classifier was employed to integrate various environmental and demographic factors, including pesticide application rates, precipitation, lithology, and population density, to predict pesticide contamination in groundwater. The models demonstrated high training accuracies of 100% and moderate average testing accuracy of 55% for atrazine and 60% for malathion across five iterations. The low test accuracy of the model, ranging from 50% to 75%, is likely due to overfitting, which can be attributed to the small dataset size and the complex nature of pesticide-contamination patterns, making it challenging for the model to generalize to unseen data. Feature importance analysis revealed that average pesticide usage emerged as the most influential factor for atrazine, while aquifer lithology and precipitation played crucial roles in both models. These results provide valuable insights into the dynamics of pesticide contamination, highlighting areas at greater risk of contamination. The findings underscore the importance of integrating environmental, geological, and agricultural variables for more effective groundwater management and sustainable agricultural practices, contributing to the protection of water resources and public health.

54 ENVIRONMENTAL SCIENCES↗

Power modeling of degraded PV systems: Case studies using a dynamically updated physical model (PV-Pro)

Power modeling, widely applied for health monitoring and power prediction, is crucial for the efficiency and reliability of Photovoltaic (PV) systems. The most common approach for power modeling uses a physical equivalent circuit model, with the core challenge being the estimation of model parameters. Traditional parameter estimation either relies on datasheet information, which does not reflect the system's current health status, especially for degraded PV systems, or requires additional I-V characterization, which is generally unavailable for large-scale PV systems. Thus, we build upon our previously developed tool, PV-Pro (originally proposed for degradation analysis), to enhance its application for power modeling of degraded PV systems. PV-Pro extracts model parameters from production data without requiring I-V characterization. This dynamic model, periodically updated, can closely capture the actual degradation status, enabling precise power modeling. PV-Pro is compared with popular power modeling techniques, including persistence, nominal physical, and various machine learning models. The results indicate that PV-Pro achieves outstanding power modeling performance, with an average nMAE of 1.4 % across four field-degraded PV systems, reducing error by 17.6 % compared to the best alternative technique. Furthermore, PV-Pro demonstrates robustness across different seasons and severities of degradation. The tool is available as a Python package at https://github.com/DuraMAT/pvpro.

14 SOLAR ENERGY↗

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (↗

Systematic characterization of unknown compounds via dimensionality reduction of time series

Analysis of ambient aerosols provides valuable insight into particle sources and formation chemistry. However, due to the complexity of atmospheric data and the dynamic nature of aerosol composition, a substantial fraction of data often become discarded by conventional analysis methods. Furthermore, a large fraction of chemical species within those data are unidentifiable due to a lack of matching spectral information, resulting in suboptimal characterization of chemical composition. Previous work has demonstrated techniques for cataloging analytes in a chromatographic dataset by deconvolution of mass spectra, but integration of these analytes throughout a large dataset remains time consuming. Here, we present a method to automatically identify an ion for quantitation for single-ion chromatogram based peak fitting and integration, enabling comprehensive integration of analytes with minimal user interaction. The resulting time series are clustered with a machine-learning based dimensionality reduction technique to systematically investigate the underlying characteristics of the categorized analytes and gain new insights into the chemical composition and physicochemical properties of the unidentifiable analytes. We apply these methods to existing atmospheric datasets collected in Manacapuru, Brazil during the GoAmazon2014/5 campaign to identify new analytes and interpret their variability and transformations in the atmosphere. The analysis results generate 408 time series from cataloged analytes of interest, and the clustering of those time series with spherical k-means results in 8 distinct clusters. We find the analytes form clusters based on their distinct physicochemical properties, demonstrating the method’s ability to systematically identify and selectively filter contaminants and instrumental analytes and characterize the unidentifiable analytes.

54 ENVIRONMENTAL SCIENCES↗

Thermal, Structural, and Rotordynamics Refinements to the Superconducting Rotor of A 1.4 Mw Partially Superconducting Machine for Electrified Aircraft

NASA is developing the high efficiency megawatt motor (HEMM), a 1.4 MW partially superconducting machine, to ad-dress the need for highly efficient, lightweight, MW-class ma-chines to support aviation sustainability efforts. This paper presents progress on the thermal, structural, and rotordynamics aspects of HEMM’s superconducting rotor. Emissivity measurements of gold coated cobalt iron and titanium parts are presented and then used to correlate a thermal model of the rotor to existing data from a thermo-electrical test of the full-scale rotor. Lessons learned from the correlated model are presented, including a rank ordering of the contributors to the temperature of the superconducting coils. An updated structural design of the HEMM rotor is presented and compared to the prior design. Finite element analysis demonstrates that the updated design improves stress margins while reducing magnetic flux leakage and retaining thermal performance. The key conclusions of the rotordynamics design are presented.

electrified aircraft propulsion↗

Machine learning-based interatomic potential development and phase transition analysis of ferroelectric hafnium dioxide

The ferroelectric phase (𝑃⁢𝑐⁢𝑎⁢2 1 , which is in orthorhombic symmetry) of hafnium dioxide (HfO 2 ) has gained much attention due to its potential applications in nanoelectronics and advanced memory devices. However, its complex phase behavior under external stimuli, such as pressure and temperature, remains a subject of intense investigation. This study focuses on developing a machine learning-based interatomic potential (MLIP) that is trained with data from density-functional theory (DFT) calculations to simulate phase transitions and mechanical properties of HfO 2 . The developed MLIP predicts lattice parameters, equations of state, bulk and shear moduli, and elastic constants that closely align with DFT predictions for several phases and at various pressures. Once validated, the MLIP is used to investigate the phase transitions of ferroelectric HfO 2 (𝑃⁢𝑐⁢𝑎⁢2 1 ) under both isobaric and constant stress conditions at elevated temperatures ranging from 200 to 2500 K. We used several complementary methods, including local symmetry identification, radial distribution function, and x-ray diffraction characterization, to identify interesting phase transitions among several competitive hafnia phases predicted from our simulations. The suggested methods uniformly reveal that under pure deviatoric condition, the system favors a transition from the orthorhombic 𝑃⁢𝑐⁢𝑎⁢2 1 phase to a tetragonal (𝑃⁢4 2 /𝑛⁢𝑚⁢𝑐) phase, whereas a zero stress condition drives the system from the 𝑃⁢𝑐⁢𝑎⁢2 1 phase to another orthorhombic (𝑃⁢𝑏⁢𝑐⁢𝑛) phase. These findings provide crucial insights into stress and temperature-induced phase behavior of hafnia, guiding future experimental and theoretical studies for optimizing hafnia-based ferroelectric devices.

Ferroelectric HfO2↗