Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

L-VISP: LSTM Visualization for Interpretable Symptom Prediction in Patient Cohorts

Symptom modelling in head and neck cancer is challenged by the complexity of heterogeneous patient data, leading to an interest in deep learning approaches. Although Long Short-Term Memory Networks (LSTMs) have shown great results in patient risk prediction, their low interpretability requires data modellers to collaborate with clinical experts to validate the results. We present L-VISP, a human–machine solution that uses visual analytics for LSTM modelling in clinical research. L-VISP uses custom visual encodings to make multiple LSTM variants interpretable, supporting a full range of analysis, from understanding model operations and evaluating performance to interpreting results in a clinical context. We evaluate L-VISP with data modellers and a clinical oncologist and present the takeaways from this multidisciplinary collaboration.

LSTM modeling↗

Fast Machine Learning Lidar Surrogate Simulator: Pristine Clear Sky

The simulations of lidar signals and retrievals rely on a range of optic-physical models, such as radiative transfer models, particle scattering and absorption models, along with the output data from atmospheric physical models. Integrating these different models to represent signals of a lidar system is computationally expensive, and performing backward retrievals can be complex and ambiguous. However, with the advantages of Machine Learning, there is a new potential for building effective lidar signal database linked to corresponding atmospheric profiles. For this project, we are developing a fast pre-trained neural network as the lidar surrogate simulator using simulated data for a CALIPSO-like lidar (355 nm, 532nm, and 1064nm), and a CO2 differential absorption lidar (DIAL) near 1571nm. Specifically, we utilize a long short-term memory (LSTM) model to map the relationships between atmospheric profiles (pressure, temperature, air density and CO2 mixing ratio) and lidar signals. This approach allows us to build machine learning based simulators that can reconstruct lidar signals at specific bands from MERRA reanalysis data, and perform retrievals of atmospheric profiles using lidar signals at various wavelengths. As a first step, the results show the potential of this method to establish a foundational model for sensor signals. This model offers the promise of enabling both accurate predictions and rapid retrievals, providing a more efficient approach to signal processing and analysis.

Shan Zeng↗

How to Build a Quantum Supercomputer: Scaling from Hundreds to Millions of Qubits

In the span of four decades, quantum computation has evolved from an intellectual curiosity to a potentially realizable technology. Today, small-scale demonstrations have become possible for quantum algorithmic primitives on hundreds of physical qubits and proof-of-principle error-correction on a single logical qubit. Nevertheless, despite significant progress and excitement, the path toward a full-stack scalable technology is largely unknown. There are significant outstanding quantum hardware, fabrication, software architecture, and algorithmic challenges that are either unresolved or overlooked. These issues could seriously undermine the arrival of utility-scale quantum computers for the foreseeable future. Here, we provide a comprehensive review of these scaling challenges. We show how the road to scaling could be paved by adopting existing semiconductor technology to build much higher-quality qubits, employing system engineering approaches, and performing distributed quantum computation within heterogeneous high-performance computing infrastructures. These opportunities for research and development could unlock certain promising applications, in particular, efficient quantum simulation/learning of quantum data generated by natural or engineered quantum systems. To estimate the true cost of such promises, we provide a detailed resource and sensitivity analysis for classically hard quantum chemistry calculations on surface-code error-corrected quantum computers given current, target, and desired hardware specifications based on superconducting qubits, accounting for a realistic distribution of errors. Furthermore, we argue that, to tackle industry-scale classical optimization and machine learning problems in a cost-effective manner, heterogeneous quantum-probabilistic computing with custom-designed accelerators should be considered as a complementary path toward scalability.

Mohseni, Masoud↗

A machine learning pipeline for identifying infiltration managed aquifer recharge locations from satellite imagery in the San Joaquin Valley, California

This study focuses on an agricultural region in California’s Central Valley, USA, where Managed Aquifer Recharge (MAR) is widely implemented to mitigate groundwater depletion under increasing water demand and climate variability. A deep learning and machine learning framework was developed to identify infiltration-MAR locations using satellite imagery and environmental data. The framework integrates surface water detection from Sentinel-2 imagery, geospatial delineation of water bodies, spatiotemporal tracking of water body dynamics, and supervised classification using meteorological, environmental, and topographic variables. The framework was applied to a 2379 km² study area southwest of Fresno, where 765 water bodies were detected, including 139 identified MAR sites based on publicly available datasets and expert knowledge. The classification model achieved an accuracy of 0.94 and an F1 score of 0.85. Feature importance analysis indicates that cropland, normalized difference vegetation index (NDVI), and evaporation are among the most influential predictors for infiltration-MAR. Notably, the framework suggests that engineered water management in infiltration-MAR systems can disrupt or even reverse the expected positive correlation between surface water extent and precipitation. These findings provide physically interpretable insights into the characteristics of existing infiltration-MAR facilities and demonstrate the potential of the proposed framework as a reproducible, interpretable, and potentially transferable tool for data-driven infiltration-MAR identification and inventory development under growing climatic and hydrological uncertainty.

Classification↗

Multiomic Network Analysis Identifies Dysregulated Neurobiological Pathways in Opioid Addiction

BACKGROUND: Opioid addiction is a worldwide public health crisis. In the United States, for example, opioids cause more drug overdose deaths than any other substance. However, opioid addiction treatments have limited efficacy, meaning that additional treatments are needed. METHODS: To help address this problem, we used network-based machine learning techniques to integrate results from genome-wide association studies of opioid use disorder and problematic prescription opioid misuse with transcriptomic, proteomic, and epigenetic data from the dorsolateral prefrontal cortex of people who died of opioid overdose and control individuals. RESULTS: Here we identified 211 highly interrelated genes identified by genome-wide association studies or dysregulation in the dorsolateral prefrontal cortex of people who died of opioid overdose that implicated the Akt, BDNF (brain-derived neurotrophic factor), and ERK (extracellular signal-regulated kinase) pathways, identifying 414 drugs targeting 48 of these opioid addiction–associated genes. Some of the identified drugs are approved to treat other substance use disorders or depression. CONCLUSIONS: Our synthesis of multiomics using a systems biology approach revealed key gene targets that could contribute to drug repurposing, genetics-informed addiction treatment, and future discovery.

60 APPLIED LIFE SCIENCES↗

Machine-Learning-Based Adaptive Thinning of CrIS Radiances to Improve Global Tropical Cyclone Analysis and Forecasts

This work is focused on optimizing the assimilation of hyperspectral infrared (IR) radiances from the Cross-track Infrared Sounder (CrIS) with the goal of improving the representation of tropical cyclones (TCs) in global analyses and forecasts. Current operational assimilation systems rely on subsampling IR radiances on a regular thinning grid. A new and improved adaptive methodology based on machine learning (ML) recognizes TCs from geostationary satellite imagery and is implemented in the Goddard Earth Observing System (GEOS) model and data assimilation framework. The ML methodology is extensively trained on existing TC data sets and creates for each TC a dynamic mask, based on the evolving shape and life cycle of that specific event. Once a TC mask is created, a switch is then activated in the data assimilation system to alter the thinning, ingesting more CrIS radiances within the moving mask, thus increasing the TC sampling. After the TC dissipates, the assimilation of CrIS radiances reverts to normal data density. Results of TC segmentation provided by a state-of-the-art generative machine learning model known as the Denoising Diffusion Probabilistic Model (DDPM) are compared to the previously used U-Net model. The new approach surpasses the performance of the previously developed one. The methodology is applied to both clear-sky and cloud-cleared radiances. Benefits from the latter methodology, particularly in improving the structure of TCs and the intensity forecasts, are presented.

Oreste Reale↗

Data from: "Towards CONUS-Wide ML-Augmented Conceptually-Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics"

This data package was generated to support the manuscript “Towards CONUS-Wide Machine Learning-Augmented Conceptually Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics.” It provides input files, model outputs, plotting data, scripts, notebooks, and documentation used to develop, evaluate, and reproduce Mass-Conserving Perceptron (MCP)-based hydrologic modeling experiments across 513 selected Catchment Attributes and Meteorology for Large-sample Studies in the United States (CAMELS-US) basins. The files are organized by modeling component and analysis purpose, including rainfall–runoff experiments, snow module experiments, coupled hydrologic-snow experiments, Long Short-Term Memory (LSTM) benchmark results, model skill metrics, initialization and epoch records, cell-state normalization files, Akaike Information Criterion (AIC)-based model comparison files, and data used to generate manuscript figures. Tabular files can be opened using standard spreadsheet software or Python/R data-analysis tools. Python scripts, Jupyter notebooks, and selected MATLAB scripts are included for model execution, postprocessing, plotting, and statistical analysis. Quality assurance and quality control were conducted through the source-data selection and modeling workflow. Meteorological forcing, streamflow, and static catchment attributes were derived from the CAMELS-US dataset, and snow water equivalent data were derived from the University of Arizona (UA) Snow Water Equivalent dataset. Selected basins and time periods were screened during the associated research workflow to avoid missing observations or poor-quality cases. Static geospatial features were processed primarily using Quantum Geographic Information System (QGIS) and Geospatial Data Abstraction Library (GDAL) workflows. Additional details are provided in the associated manuscript and documentation.

ESS-DIVE CSV File Formatting Guidelines Reporting ↗

Soil Salinity Level Assessment and Prediction Integrating UAV-borne Hyperspectral Imaging and Machine Learning Algorithms to Combat Desertification

In response to the ongoing global food crisis, the United Nations has identified “Zero Hunger” as one of its Sustainable Development Goals. A central contributor to the crisis is the process in which agricultural lands go through desertification. Research has shown a direct correlation between soil salinity and desertification - increased salinity levels indicate a higher risk for desertification. Furthermore, researchers have explored various techniques to map soil salinity, but these methods are oftentimes inefficient and don’t address future salinity predictions. To improve desertification monitoring, soil salinity can be observed via hyperspectral imaging on unmanned aerial vehicles (UAVs) to predict the risk of agricultural desertification using artificial intelligence (AI) and machine learning (ML) techniques. A significant gap exists in past research that applies ML and imaging techniques to soil salinity: convolutional neural networks (CNNs) and regression models are rarely leveraged together, despite the efficiency and accuracy of these models. To compensate for this gap, the proposed system leverages the use of these AI and ML models to improve soil assessment and prediction techniques. This approach involves three steps - data collection, image analysis, and future prediction. Using hyperspectral cameras on UAVs to collect the data from the region, a trained CNN model will output estimated soil salinity levels at a specific time. The estimations will then be analyzed by a regression model to assess the accuracy of future soil salinity predictions. The proposed system will identify regions at risk of desertification to help farmers mitigate agricultural loss, in turn helping alleviate the food crisis.

UAV systems↗

Using convolutional neural networks to detect edge localized modes in DIII-D from Doppler backscattering measurements

In H-mode tokamak plasmas, the plasma is sometimes ejected beyond the edge transport barrier. These events are known as edge localized modes (ELMs). ELMs cause a loss of energy and damage the vessel walls. Understanding the physics of ELMs, and by extension, how to detect and mitigate them, is an important challenge. In this paper, we focus on two diagnostic methods—deuterium-alpha (D α ) spectroscopy and Doppler backscattering (DBS). The former detects ELMs by measuring Balmer alpha emission, while the latter uses microwave radiation to probe the plasma. DBS has the advantages of having a higher temporal resolution and robustness to damage. These advantages of DBS diagnostic may be beneficial for future operational tokamaks, and thus, data processing techniques for DBS should be developed in preparation. In sight of this, we explore the training of neural networks to detect ELMs from DBS data, using D α data as the ground truth. With shots found in the DIII-D database, the model is trained to classify each time step based on the occurrence of an ELM event. The results are promising. When tested on shots similar to those used for training, the model is capable of consistently achieving a high f1-score of 0.93. Furthermore, this score is a performance metric for imbalanced datasets that ranges between 0 and 1. We evaluate the performance of our neural network on a variety of ELMs in different high confinement regimes (grassy ELM, RMP mitigated, and wide-pedestal), finding broad applicability. Beyond ELMs, our work demonstrates the wider feasibility of applying neural networks to data from DBS diagnostic.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Cluster Analysis of Spectroscopic Line Profiles and EUV Emission in RMHD Simulations and Observations of the Solar Atmosphere

Spatially-resolved observations from the IRIS, SDO/AIA, and other space mission and ground-based telescopes, coupled with realistic 3D RMHD simulations, are a powerful tool for analysis of processes in the solar atmosphere. To better understand the dynamical and thermodynamic properties in the simulation data and their connection to observations, it is essential to determine similarities in the behaviors of the synthesized and observed emission. However, the complexity of observational data and physical processes makes comparison of observations and modeling results difficult. In this work, we show the initial results of application of K-Means clustering (unsupervised machine learning) algorithm to two different problems: 1) recognition of the typical spectroscopic line profiles observed by IRIS during solar flares and their typical dynamic behavior; 2) recognition of shocks and heating events in synthetic AIA emission data obtained from StellarBox quiet-Sun simulations. The average silhouette width technique for the KMeans algorithm is utilized in different ways to obtain optimal numbers of clusters. We discuss application of the emission clustering to visualizations of the computational volume, understanding its evolutionary trends and behavior patterns, and inversion (reconstruction) of physical properties of the solar atmosphere from synthesizes emission data.

Sadykov, Viacheslav↗

Integration of LIBS with Machine Learning for Real-Time Monitoring of Feedstock in H 2 Gasification Applications

This project, funded by the U.S. Department of Energy (DOE) – Office of Fossil Energy under Award Number DE-FE0032177, aimed to assess the feasibility of an integrated Laser-Induced Breakdown Spectroscopy (LIBS) system with advanced machine learning (ML) models for real-time characterization and potential control of hydrogen gasifiers running on waste materials as feedstocks. This was a multidisciplinary effort that encompassed the acquisition and standardized analysis of individual and blended feedstocks—comprising biomass, coal waste, and plastic waste, followed by the development of a dynamic LIBS bench system for material sample analysis and development of predictive ML models. Comprehensive laboratory testing enabled the creation of a robust elemental dataset that served as the foundation for ML model training. Techniques such as Random Forest, Gradient Boosting, Support Vector Regression, and Neural Networks were employed to predict key feedstock properties, including higher heating value (HHV), moisture content, thermal conductivity, and ash composition with high accuracy. The results were validated against experimental data and demonstrated strong potential for real-time application in gasifier control systems. The project concluded with a study on the integration of the LIBS+ML approach for gasifier control and a techno-economic analysis of the implementation of the approach into hydrogen (H 2 ) gasification systems. Dissemination of results was carried out at a DOE meeting. This work establishes a scalable framework for automated, in-line feedstock quality assessment, offering significant implications for process optimization and emissions reduction in hydrogen production.

01 COAL, LIGNITE, AND PEAT↗

Dynamic STEM-EELS for single-atom and defect measurement during electron beam transformations

This study introduces the integration of dynamic computer vision–enabled imaging with electron energy loss spectroscopy (EELS) in scanning transmission electron microscopy (STEM). This approach involves real-time discovery and analysis of atomic structures as they form, allowing us to observe the evolution of material properties at the atomic level, capturing transient states traditional techniques often miss. Rapid object detection and action system enhances the efficiency and accuracy of STEM-EELS by autonomously identifying and targeting only areas of interest. This machine learning (ML)–based approach differs from classical ML in that it must be executed on the fly, not using static data. We apply this technology to V-doped MoS 2 , uncovering insights into defect formation and evolution under electron beam exposure. This approach opens uncharted avenues for exploring and characterizing materials in dynamic states, offering a pathway to increase our understanding of dynamic phenomena in materials under thermal, chemical, and beam stimuli.

47 OTHER INSTRUMENTATION↗

Mapping National Forest Aboveground Biomass in Mexico by Integrating GEDI, Sentinel‐1 and Sentinel‐2 Data

Accurate mapping of forest aboveground biomass density (AGBD) is required to better understand the role of forests in the global carbon cycle and to support international policies for climate change mitigation and adaptation. Mexico is one of the countries having great potential for the United Nations Programme on Reducing Emissions from Deforestation and Forest Degradation (or UN-REDD program) and there is a growing demand for unbiased Monitoring Reporting Verification systems at a national level. As an effort under NASA’s Carbon Monitoring System (CMS) program, we developed a machine learning model using multi-stream remote sensing measurements as well as topographic data to create a high spatial resolution AGBD map (~100 m) over Mexico (circa 2020). The remote sensing data includes Global Ecosystem Dynamic Investigation (GEDI) lidar, Sentinel 1 Synthetic-Aperture Radar (SAR), and Sentinel-2 multispectral imagery (MSI). GEDI onboard the International Space Station provides unprecedented forest structure and AGBD sampling datasets for model training and validation practices. Our analysis indicates that the developed random forest model can capture 63 % of the spatial variation (RMSE = 33.7 Mg/ha) of AGBD of Mexican forests. We find that shortwave infrared bands of Sentinel-2 MSI and topographical variables from elevation data are the most important variables in the developed AGBD model. Our study highlights methodological opportunities in synergistic uses of multiple sensors for large-scale forest AGBD mapping and shows potential for retrospective analysis and operational monitoring of forest AGBD and its dynamics.

Taejin Park↗

Machine Learning Approaches to Data Reduction from the MapX X-ray Fluorescence Instrument for Detection of Biosignatures and Habitable Planetary Environments

The search for evidence of life or its processes involves the detection of biosignatures suggestive of extinct or extant life, or the determination that an environment either has or once had the potential to harbor life. In situ elemental imaging is useful in either case, since features on the mm to μm scale reveal geological processes which may indicate past or present habitability. The Mapping X-ray Fluorescence Spectrometer (MapX) is an in-situ instrument designed to identify these features on planetary surfaces. Here we present progress on instrument development, data analysis methods, and element quantification.

Walroth, Richard C.↗

From observation to replication: machine-learning-driven quantification and replication of fine-scale fish kinematics and behavior

Long-term quantification of fish behavior is essential for aquatic ecology, wildlife telemetry, and biomechanical device development. However, the observation duration required to obtain reliable behavioral and kinematic metrics remains unclear, and few tools exist to physically reproduce natural swimming motion for controlled experimentation. We address these challenges by developing a generalizable framework that models behavioral reliability (Spearman–Brown reliability index) as a function of observation duration and derives metric-specific monitoring thresholds. Using juvenile white sturgeon (Acipenser transmontanus) as a case study, we demonstrate that the minimum duration needed for reliable estimates varies substantially across kinematic features: to exceed a reliability of 0.8, total distance traveled requires 12 days, average curvature (mm?¹) 15 days, tail-beat frequency (Hz) 8 days, and average speed (body length/s) 17 days. We further bridge digital analysis and physical testing by developing a hardware-in-the-loop simulator that reconstructs machine-learning-derived swimming kinematics with high fidelity (correlation coefficient 0.98–0.99, RMSE 1.22–1.27 mm over a 5-minute segment). This platform enables realistic, repeatable motion stimuli for evaluating aquatic sensing technologies and bio-integrated devices under controlled conditions. Together, these contributions provide a scalable approach for designing long-term behavioral studies and a data-driven connection between ecological observation and robotic experimentation.

Hwang, SungJoo↗

Gaining the most utility from our geospace observational system: Network analysis of total electron content as a means to understand space weather to the point of prediction

We present the first network analysis of interplanetary magnetic field (IMF) clock angle dependent, high-latitude, hemispheric-specific total electron content (TEC) data. We examine network parameters to describe spatio-temporal correlations in the TEC data for January 2016. We find that significant network structure exists distinguishing the dayside and nightside ionosphere, and specific features in the high-latitudes (cusp/ionospheric footpoints of magnetospheric boundary layers, polar cap, and auroral zone), and that these features vary with IMF clock angle. In this brief summary paper, we provide proof of concept results and identify important areas of future research, providing a basis for the discussion of network analysis and machine learning approaches for space weather applications.

Malik, Nishant↗

Benchmarking machine learning interatomic potentials via phonon anharmonicity

Abstract Machine learning approaches have recently emerged as powerful tools to probe structure-property relationships in crystals and molecules. Specifically, machine learning interatomic potentials (MLIPs) can accurately reproduce first-principles data at a cost similar to that of conventional interatomic potential approaches. While MLIPs have been extensively tested across various classes of materials and molecules, a clear characterization of the anharmonic terms encoded in the MLIPs is lacking. Here, we benchmark popular MLIPs using the anharmonic vibrational Hamiltonian of ThO 2 in the fluorite crystal structure, which was constructed from density functional theory (DFT) using our highly accurate and efficient irreducible derivative methods. The anharmonic Hamiltonian was used to generate molecular dynamics (MD) trajectories, which were used to train three classes of MLIPs: Gaussian approximation potentials, artificial neural networks (ANN), and graph neural networks (GNN). The results were assessed by directly comparing phonons and their interactions, as well as phonon linewidths, phonon lineshifts, and thermal conductivity. The models were also trained on a DFT MD dataset, demonstrating good agreement up to fifth-order for the ANN and GNN. Our analysis demonstrates that MLIPs have great potential for accurately characterizing anharmonicity in materials systems at a fraction of the cost of conventional first principles-based approaches.

interatomic potentials↗