Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

TransPlatformer

We propose TransPlatformer for translating toxicogenomics from one platform to another. Transcriptomic profiling has evolved through multiple generations of technology, from microarrays (e.g., Affymetrix, CodeLink) to more recent high-throughput sequencing and targeted panels such as S1500+. Microarrays, which dominated gene expression studies in the early 2000s, provided affordable and high-throughput transcript quantification but suffered from cross-hybridization issues and limited dynamic range . RNA-Seq, introduced in the late 2000s, revolutionized transcriptomics by enabling unbiased and comprehensive gene expression analysis, albeit at higher costs and computational demands . Despite advances, many studies rely on historical microarray data, necessitating the translation of legacy data into modern platforms to ensure continuity and comparability. This translation is complicated by factors such as platform-specific probe design, differences in transcript coverage, and batch effects . Existing methods for cross-platform mapping include statistical normalization, machine learning models, and biological anchoring approaches. The ability to translate transcriptomic data between platforms has broad implications, including enhanced meta-analyses, improved toxicological modeling, and better integration of historical datasets with contemporary research. TransPlatformer seeks to contribute to this effort by evaluating translation methodologies and proposing novel strategies to improve cross-platform gene expression harmonization. In this repository there are code examples for TransPlatformer implementation

Cong, Guojing↗

Landscaper v1

Understanding the inner workings of machine learning models through their loss landscapes offers crucial insights into model properties, optimization dynamics, and generalizability. However, accessing these insights has traditionally required specialized mathematical expertise, limiting broader adoption. Landscaper is an open-source Python package designed to bridge this gap. Landscaper seamlessly integrates a suite of multi-dimensional loss landscape analyses with cutting-edge topological data analysis (TDA) methods. This powerful combination makes both fundamental loss landscape analysis and advanced TDA techniques accessible to the broader scientific ML community, without requiring deep pre-existing mathematical knowledge. Landscaper offers three key functionalities: * Construction: Builds detailed loss landscape representations through versatile low and high-dimensional sampling techniques. * Quantification: Applies advanced metrics, including a novel topological data analysis (TDA) based smoothness metric, enabling new perspectives on model behavior. * Visualization: Offers intuitive tools to visualize and interpret loss landscapes, providing actionable insights beyond traditional performance metrics.

Weber, Gunther [Lawrence Berkeley National Laborat↗

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference↗

SSTDR and FDR Detection of Un-Energized and Energized Cable Anomalies Including Thermal Degradation Using Machine Learning

Historically, cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, continued use of these cables must shift to a performance-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. A variety of cable tests are available and are commonly applied during outages when the cables can be taken out of service. Frequency domain reflectometry (FDR) is one of these test methods that is being more broadly accepted and used because it not only detects anomalies along the cable with a low-voltage signal that does not stress the cable insulation, but the technique also locates the anomalies. This supports follow-up local inspection and local repair or partial replacement of a damaged cable segment. Currently, FDR testing is only applied to cables that are taken out of service since the test instrument would be damaged by operational voltages. A related technology that has found some acceptance in the aircraft and rail industry is spread spectrum time domain reflectometry (SSTDR). This technology has been implemented with a custom commercial instrument by LiveWire Innovation that is designed to operate on live cables up to 1000 volts and with a bandwidth of 48 MHz. Initial evaluation by the Pacific Northwest National Laboratory (PNNL) of the Live Wire system indicated that a broader bandwidth (BW) SSTDR may be better for many kinds of flaws. This led PNNL to develop an SSTDR laboratory instrument suitable for tests up to 500 MHz bandwidth. Testing on energized cables is also desirable for online monitoring systems so an inductive clamshell coupler was developed that allows energized cables to be tested up to at least 5 kV and likely higher voltage levels. Dielectric spectroscopy and tan delta testing plus various laboratory destructive tests were included in this data acquisition campaign directed to feed a machine learning (ML) study. With these kinds of developments, online energized cable tests may be possible with industrial adoption of such hardware advances but it will be completely impractical to have highly skilled data analysts continually examine these complex signals for indications of damage or compromised conditions. If online testing is to be implemented in new test hardware, it must be accompanied by software that can interpret the signals and alert plant operators of changing or degraded conditions. The thermally aged, shielded cable investigated here was separately treated for ML analysis. Visual analysis of electrical data showed generally increasing peaks where the cable entered and exited the oven. These peaks were not exactly aligned with expected locations, but these differences were attributed to velocity of propagation calibration errors. Only supervised ML was applied to the thermally aged data as this data was only available shortly before the committed publication date of this report. The supervised ML was structured to divide the 0 to 70-day responses as ‘normal’ from 0 to 35 days or ‘anomalous’ from 36 to 70 days, based on cable tensile elongation at break (EAB) insulation characterization. Using 80% of the data for training and 20% for testing, the supervised ML predicted normal versus anomalous was 70% accurate. Important conclusions include: • Accuracy to predict the presence of cable damage is improved from the 2023 effort by more training data. Weighted accuracies for comparisons among the instruments ranged from 67 to 89 % for unsupervised ML and 71 to 99% for supervised ML. • Based on the synthetic data tests, the unsupervised models are more generalizable to unseen anomalies. The Multi-Layer Perceptron classifier (MLP) model reported as high as 99.7% accuracy on the test data, but this dropped to 58.3% when tested on the synthetic data. In contrast, the unsupervised Pointwise model only achieved 89.7% accuracy on the experimental data but reported 78.3% accuracy on the synthetic data. • The best anomaly indicators are higher frequency (400 MHz BW) FDR data. Other tests may be interesting but for this study, this was the best predicter.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A structured framework for predicting sustainable aviation fuel properties using liquid-phase FTIR and machine learning

Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.

Fourier transform infrared spectroscopy↗

ASCoT 3: Nonlinear Principal Components Analysis and Uncertainty Quantification in Early Concept Spacecraft Flight Software Cost Estimation

For mission planners and evaluators alike, value in cost models comes from a mean or median prediction, an understanding of the uncertainty on that prediction, and an understanding of model performance. Here we apply advanced statistical and machine learning methods to spacecraft flight software cost, effort, and SLOC estimation, and present the results in the latest version of the Analogy Software Cost Tool (ASCoT). We present in- and out-of-sample performance metrics for our models, each of which incorporate some amount of epistemic uncertainty. ASCoT, hosted on the One NASA Cost Engineering (ONCE) database via the Online NASA Space Estimation Tool (ONSET), was first showcased in 2016 as a number of analogy-based models and methods (kNN and Clustering) to support early project formulation. This ASCoT update improves upon the previous analogic methods by incorporating uncertainty in the data transformations. In particular, we use a Nonlinear Principal Components Analysis (NLPCA) to deal with ordinal data.

Robotic Spacecraft↗

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING↗

Integrating Maximum Entropy Production Theory and Machine Learning to Improve Global Evapotranspiration Modeling

Accurate estimation of terrestrial evapotranspiration (ET) is vital for understanding global water and energy cycles. However, current global ET estimations are not well constrained. This study introduces an integrated framework combining the Maximum Entropy Production (MEP) theory with Random Forest (RF) model to improve global ET estimation. Specifically, in contrast to direct ET estimation by the RF model, the integrated framework (MEP‐RF) trains to predict error of MEP‐simulated ET. MEP‐RF outperforms RF in spatiotemporal extrapolation. Attribution analysis with in situ observations reveals that the inputs of MEP are the most critical variables for the ET process, including net radiation, vegetated area, soil moisture, and surface temperature. We further drive MEP‐RF with global reanalysis and satellite data sets of these four inputs, yielding a global mean terrestrial ET of 548 mm/year, with 77% attributed to transpiration. The global ET increased at a rate of 0.85 mm/year per year during 2003–2021, primarily due to vegetation greening rather than rising temperature, while decreasing soil moisture led to decreasing regional ET. The integrated framework provides a novel approach for the estimation of global ET without the need for hard‐to‐obtain and thus uncertain inputs, such as wind speed, surface roughness, aerodynamic and canopy stomatal resistance. Therefore, MEP‐RF offers an independent method on existing global ET products. It represents a promising physically based approach that can be incorporated into Earth System Models to enhance water and energy cycle simulations.

54 ENVIRONMENTAL SCIENCES↗

Predicting non-linear stress–strain response of mesostructured cellular materials using supervised autoencoder

Recent breakthroughs in advanced manufacturing capabilities have made it possible to design and print sophisticated topologies of cellular structures using diverse engineering materials such as metals, polymers, and ceramics. In these architectured materials, it is often desirable to tailor the mechanical properties by altering the unit cell topology. This necessitates an in-depth understanding of how the topology of the unit cell structure affects the macroscopic behavior of the material in both the linear and the non-linear regimes encountered under large compression. Here, we have developed a machine learning (ML) approach capable of accelerating the prediction of the stress–strain response of a polymer-based cellular structure under uniaxial confined compression. As part of generating the training data for ML, 60,000 mesostructures were generated using a relatively novel approach based on cellular automata, and their corresponding stress–strain responses were obtained from the finite element simulations. Principal component analysis (PCA) was used to reduce the dimensionality of the stress–strain curves. With only 20 principal components, PCA captured 99.89% of the variance in the stress–strain curves while reducing the dimensionality by 5X. ML using supervised autoencoder was able to successfully speed up the prediction of the non-linear stress–strain response of a unit cell by up to 4600X. The proposed method can serve as an efficient data generation tool and a rapid means for predicting the structure–property relationship through accelerated forward modeling of cellular materials under compaction, in cases where the macroscopic stress–strain response is governed by the unit-cell topology.

36 MATERIALS SCIENCE↗

Unsupervised Learning for Improved Gamma-Ray Spectrometry in Pixelated Cadmium Zinc Telluride (CZT) Detectors

Machine learning has been found to be ubiquitously useful across many industries, presenting an opportunity to improve radiation detection performance using data-driven algorithms. Improved detector resolution can aid in the detection, identification, and quantification of radionuclides. Here, in this work, a novel, data-driven, unsupervised learning approach is developed to improve detector spectral characteristics by learning, and subsequently rejecting, poorly performing regions of the pixelated detector. Feature engineering is used to fit individual characteristic photo peaks to a Doniach lineshape with a linear background model. Then, principal component analysis is used to learn a lower-dimension latent space representation of each photo peak where the pixels are clustered, and subsequently ranked, based on the cluster mean distance to an optimal point. Pixels within the worst cluster(s) are rejected to improve the full-width at half-maximum (FWHM) by 10% to 15% (relative to the bulk detector) at 50% net efficiency when applied to training data obtained from measurements of a 100 μCi 154 Eu source using a H3D M400i pixelated cadmium zinc telluride detector. These results compare well with, but do not outperform, a greedy algorithm that accumulates pixels in order of FWHM from lowest to highest used as a benchmark. In the future, this approach can be extended to include the detector energy and angular response. Finally, the model is applied to newly seen natural and enriched uranium spectra relevant for nuclear safeguards applications.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Analysis of Slow Spill Data for the Mu2e Experiment

The execution of the Mu2e experiment requires a stable, low-intensity proton beam from the Delivery Ring to produce clean data and protect equipment. This is done by performing a “slow extraction,” which is the gradual contraction of the stable region within the accelerator’s beam pipe. The Delivery Ring is currently unable to perform slow extraction with the stability required by Mu2e. To resolve this, the FAN-C team is training machine learning models with the purpose of replacing the Delivery Ring’s current PID controllers with AI-powered controllers. Training these models requires clean, processed data from slow spills. Over the course of this project, data from previous slow spills were processed and analyzed, and the clean data, graphs, and insights gained from the process were provided to the FAN-C team to assist them in their efforts.

Osborn, Thomas [Purdue U., West Lafayette]↗

Mic-hackathon 2024: hackathon on machine learning for electron and scanning probe microscopy

Microscopy is one of the primary sources of information on materials structure and functionality at the nanometer and atomic scales. The data generated through microscopy is often contained in well-structured datasets, enriched with extensive metadata and sample histories, although not always with the same level of detail or storage format. The broad incorporation of data management plans by major funding agencies ensures the preservation and accessibility of this data. However, deriving insights from these rich datasets remains challenging due to the lack of established code ecosystems, standardized benchmarks, and integration strategies. Correspondingly, the efficiency of data usage is very low, and time expenditures at the analysis stage are enormous. In addition to post-acquisition data analysis, the emergence of application programming interfaces by major microscope manufacturers now creates opportunities for real-time ML-based data analytics to enable automated decision making, and particularly ML-agent controlled real-time microscope operation. Despite these opportunities, there is a significant gap in integrating the ML community with the broader microscopy community, limiting the value that these methods bring to physics and materials discovery and materials optimization. Hackathons address these challenges by fostering collaboration between ML experts and microscopy professionals, encouraging the development of innovative solutions that leverage ML for microscopy and preparing the workforce of the future both for microscopy-intensive domains areas, instrument manufacturers, and ML scientists interested in real world applications for fundamental research, materials optimization, and manufacturing. The hackathon generated benchmark datasets and digital twins of microscopes that further contribute to the development of the field and establish data analysis ecosystems. All the codes can be found at GitHub(https://github.com/KalininGroup/Mic-hackathon-2024-codes-publication/tree/1.0.0.1) and Zenodo (https://zenodo.org/records/15579940).

97 MATHEMATICS AND COMPUTING↗

A Low-Power High-Speed Smart Sensor Design for Space Exploration Missions

A low-power high-speed smart sensor system based on a large format active pixel sensor (APS) integrated with a programmable neural processor for space exploration missions is presented. The concept of building an advanced smart sensing system is demonstrated by a system-level microchip design that is composed with an APS sensor, a programmable neural processor, and an embedded microprocessor in a SOI CMOS technology. This ultra-fast smart sensor system-on-a-chip design mimics what is inherent in biological vision systems. Moreover, it is programmable and capable of performing ultra-fast machine vision processing in all levels such as image acquisition, image fusion, image analysis, scene interpretation, and control functions. The system provides about one tera-operation-per-second computing power which is a two order-of-magnitude increase over that of state-of-the-art microcomputers. Its high performance is due to massively parallel computing structures, high data throughput rates, fast learning capabilities, and advanced VLSI system-on-a-chip implementation.

Fang, Wai-Chi↗

EVA Task and 3D Pose Recognition from Video

Extravehicular Activity (EVA) has been known to involve potential risks of biomechanical stresses and injuries to crewmembers. Gathering of EVA motion patterns is necessary for risk analysis and mitigation. However, many existing techniques, such as motion capture systems, are not only cost-prohibitive but are impractical for retrospective analysis of past missions. In this work, a software tool was developed, which can estimate the 3D poses of a spacesuit from photographs or videos, without using special sensors or equipment. The tool is based on the state-of-the-art artificial intelligence and machine learning (AI/ML) system, which was trained by studying and capturing motion patterns of past and current spacesuit test data. The AI/ML tool was further enhanced using synthetically generated data, in which the suit postures, backgrounds, camera angles and illumination conditions were parametrically adjusted and rendered for training. The tool, incorporated the methodologies of Convolutional Neural Network (CNN), was trained, and tested in the cloud computing environment. The trained model was then applied on new imagery and video to extract estimated joint positions and suit outlines. The joint positions were further processed to capture activity (“digging”), pose labels (“bending”), and other useful downstream information. The model performance on new imagery and video was successfully assessed for accuracy and reliability. This AI/ML based posture recognition tool thus allows for the quantification of injury risk and task performance characterization for both current and past missions and training, which can immensely help to improve EVA task and suit design.

Kyung Han Kim↗

Toward first principles-based simulations of dense hydrogen

Accurate knowledge of the properties of hydrogen at high compression is crucial for astrophysics (e.g., planetary and stellar interiors, brown dwarfs, atmosphere of compact stars) and laboratory experiments, including inertial confinement fusion. There exists experimental data for the equation of state, conductivity, and Thomson scattering spectra. However, the analysis of the measurements at extreme pressures and temperatures typically involves additional model assumptions, which makes it difficult to assess the accuracy of the experimental data rigorously. On the other hand, theory and modeling have produced extensive collections of data. They originate from a very large variety of models and simulations including path integral Monte Carlo (PIMC) simulations, density functional theory (DFT), chemical models, machine-learned models, and combinations thereof. At the same time, each of these methods has fundamental limitations (fermion sign problem in PIMC, approximate exchange–correlation functionals of DFT, inconsistent interaction energy contributions in chemical models, etc.), so for some parameter ranges accurate predictions are difficult. Recently, a number of breakthroughs in first principles PIMC as well as in DFT simulations were achieved which are discussed in this review. Here we use these results to benchmark different simulation methods. We present an update of the hydrogen phase diagram at high pressures, the expected phase transitions, and thermodynamic properties including the equation of state and momentum distribution. Furthermore, we discuss available dynamic results for warm dense hydrogen, including the conductivity, dynamic structure factor, plasmon dispersion, imaginary-time structure, and density response functions. We conclude by outlining strategies to combine different simulations to achieve accurate theoretical predictions that are based on first principles.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Fracture Network Prediction Using Physics-based Machine Learning Algorithms

In recent years, systematic CO2 injection into geological reservoirs across the U.S. has gained traction as a strategy to mitigate greenhouse gas emissions. This approach necessitates precise monitoring to ensure secure containment, minimize risks, and optimize storage management. Our study leverages machine learning (ML) techniques to advance the understanding of CO2 injection processes, focusing on the Illinois Basin. Over a three-year injection period, we analyzed microseismic data, identifying 19 temporal intervals with significant bottom-hole pressure changes. By partitioning microseismic events into these intervals and estimating b-values, we revealed over 100 clusters of events related to fracture initiation or reactivation. Advanced spatial analysis highlighted horizontally-oriented fractures along the NNW-SSE axis. This quantification of fracture networks informs dynamic injection scheduling, work-over strategies, and risk assessments, enhancing carbon capture, utilization, and storage (CCUS) operations. Additionally, our methodology offers valuable insights for oil and gas operations and geothermal development, supporting fracture-based monitoring and risk mitigation.

Kumar, Abhash↗

Automating Detection and Diagnosis of Faults, Failures, and Underperformance in PV Plants

The project developed hybrid physics-based and machine-learning methods for near-real-time detection of balance-of-system faults (e.g., string, combiner, and tracker outages) in utility-scale Photovoltaic plants, achieving over 50% true positive rates with under 10% false positives and significantly reducing engineering setup time. In the extended phase, the scope expanded to plant-level underperformance analysis and industry benchmarking through the SUPER.epri.com platform. SUPER standardizes data processing and performance metrics across more than 9 GWac and 120+ plants, enabling robust comparisons and insights into loss rates, inverter downtime, and capacity degradation.

14 SOLAR ENERGY↗

Global Evaluation of Process Conditions and Wave Modes in a Rotating Detonation Engine

Rotating Detonation Engines (RDEs) show significant promise for enhancing the efficiency of gas turbine engines while maintaining low 𝑁𝑂𝑥 emissions. This work investigates the predictability of wave modes in a water-cooled RDE under varying operational conditions. Experimental data comprising over 6,700 samples was collected, including parameters such as flow rates, temperatures, pressures, and equivalence ratios. A machine learning approach using the XGBoost library was used to build a multi-class classifier, predicting wave modes based on these inputs. The model achieved a high accuracy of 97%, demonstrating that wave modes are not random but deterministic based on the process conditions. SHAP analysis was used to identify the most influential parameters affecting wave mode prediction. The results show that for the water-cooled NETL RDE, wave mode is determinant and predictable based on the process parameters.

Weber, Justin [NETL] (ORCID:0000000218487035)↗