Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine learning models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Complementing the CCS Class VI Well Permit Process with DOE-NETL's SMART Initiative Tools and Workflows

This is a presentation on model explorer developed under SMART initiative Task 2. Our team will present the current progress of the model explorer in using machine learning models to accelerate CCS project at GWPC meeting. Model explorer bring new capabilities, (fast, Realtime, and accurate) that can help CCS stakeholders including regulatory agencies, public and site operators make faster decisions and process information and data.

Hosseini, Seyyed↗

An Overview of Ecological Modeling and Machine Learning Research Within the U.S. National Aeronautics and Space Administration

In the early 1980 s NASA began research to understand global habitability and quantify the processes and fluxes between the Earth's vegetation and the biosphere. This effort evolved into the Earth Observing System Program which current encompasses 18 platforms and 80 sensors. During this time, the global environmental research community has evolved from a data poor to a data rich research area and is challenged to provide timely use of these new data. This talk will outline some of the data mining research NASA has funded in support for the environmental sciences in the Intelligent Systems project and will give a specific example in ecological forecasting, predicting the land surface properties given nowcasts and weather forecasts, using the Terrestrial Observation and Prediction System (TOPS).

Coughlan, Joseph C.↗

Towards a program of record of inland water quality: Exploiting present and heritage multispectral sensors for maximum information extraction

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This research exploits recent advancements in bio-optical modeling, cloud computing, and machine learning to enhance our capacity to leverage present and heritage satellite data. Recent research suggests that sensors with low spectral resolution, such as Sentinel 2 and Landsat missions, contain enough hidden spectral variation which can be exploited using data-driven approaches. The availability of three decades of archival imagery will open doors to discover global trends of eutrophication and increased cyanobacteria dominance and provide valuable insight to the development of predictive methodologies. Preliminary efforts in synthetic emulation of global natural inland waters will be discussed and contextualized against satellite radiometric measurement uncertainty, satellite data product uncertainties and causal signal ambiguity over the visible wavelength range, supported by high quality field and image data for selected inland aquatic sites. Insights on water quality estimation via data-driven machine learning models versus matrix inversions will be discussed, and how we can exploit spectral-spatial relationships in high spatial resolution data. A cross-sensor synergistic approach with detailed uncertainty analysis based on optical water types, will allow for unprecedented global snapshots of fine scale ecological dynamics of inland waters.

Inland↗

Development and Application of NASA SPoRT’s DustTracker-AI Model for Real-Time Identification and Tracking of Dust in Geostationary Satellite Imagery

The NASA Short-term Prediction Research and Transition (SPoRT) Center developed the DustTracker-AI model for identifying and tracking dust in NASA/NOAA Geostationary Operational Environmental Satellite (GOES) imagery in a real-time framework. A training dataset consisting of day and night dust cases was gathered over the southwestern consisting of 115 distinct images and over a million dust pixels and 256 million no dust pixels. The dataset was separated into training (60%), testing (20%), and validation (20%). A simple random forest machine learning model was developed originally to overcome the problem of night-time dust detection and has been expanded to a comprehensive day/night model for dust identification and tracking. This physically-based machine-learning approach uses NASA/NOAA GOES-16 Advanced Baseline Imager infrared imagery as inputs to the model. The model probability of dust output achieves an Area-Under-Curve (AUC) of 0.97 with a standard deviation of 0.04 for dust cases. For images with dust present, the model correctly labels 85% of dust pixels for all dust images in the validation data set. In conjunction with developing the machine-learning model, the NASA Short-term Prediction Research and Transition Center (SPoRT) partnered with NOAA National Weather Service forecast offices to evaluate the model for utility in weather forecasting operations during the 2021 and 2023 late winter-spring seasons. Preliminary evaluation has indicated the majority of forecasters described the DustTracker-AI probabilities as having added confidence to interpreting the Dust RGB and other satellite products to objectively assess the dust extent and trends and increased the amount of time the dust plume could be tracked into the night as compared to use of the Dust RGB. More recently, SPoRT tested small scale events associated with thunderstorm outflow and burn scars to determine the model’s ability to capture local events. This presentation highlights design of the model, validation/evaluation of model performance, and example cases collected during end user product assessments.

Connor H Welch↗

MLSPICE: Machine Learning based SPICE Modeling Platform for Power Magnetics

Electrical power converters are critical to a wide range of applications ranging from renewable integration to transportation electrification, and can be a key factor determining the size, weight, and efficiency of energy conversion systems. Magnetic components are typically the largest and least efficient components in power electronics. While there have been major strides in the modeling and analysis of power semiconductor devices and circuit simulations, the necessary advances in the design of power magnetics have lagged. In this project, we have transformed the modeling and design of power magnetics with machine learning enabled methods and catalyze simultaneous disruptive improvements for ML-based power electronics design tools. A fully automated open-source machine learning based magnetics modeling platform – the MagNet project - with innovations in full stack have been developed to greatly accelerate the design process and provide new insights to magnetic material and geometry design. The ARPA-E funded MagNet platform contains three major building blocks: 1) a ML-Integrated Data Acquisition System (MIDAS): a highly automated data acquisition testbed which is capable of measuring a large number of magnetic cores with a wide range of electrical circuit excitations; 2) a ML-integrated Core Loss Model (MICLM): a machine-learning trained modeling method for modeling the core loss and saturation effects of magnetic materials for arbitrary excitation waveforms; 3) ML-guided Magnetics SPICE Simulation Tool (PMSPICE): a fully integrated CAD tool which can simulate the magnetics in SPICE. It can help the designers to quickly model the linear and non-linear characteristics of magnetic components and evaluate their behavior in SPICE simulations. The developed MagNet system has fully demonstrated the proposed performance target and has been open sourced to the entire power electronics community to advance the modeling and design of power magnetics from many different angles.

36 MATERIALS SCIENCE↗

Machine Learning the COSMO Model for Predicting Thermodynamics of Electrolyte Mixtures

Bottom-up design of electrolyte mixtures for battery systems requires predicting macro thermodynamic properties from molecular constituents. For instance, molten salt electrolyte batteries require conditions far above room temperature to operate. Therefore, discovering mixtures with increasingly lower eutectic melting points is desirable. A model that can approximate chemical activity is a valuable tool to search through the vast compositional design space. Machine learning can predict properties of materials such as vibrational free energies, electronic energy gaps, and thermal conductivities. Moreover, they can learn physical models such as interatomic potentials. The COSMO-SAC model uses theory and empirical parameterization to predict liquid-vapor and liquid-solid properties using first-principles calculations. However, obtaining activity coefficients required for parameterizing the COSMO-SAC model is costly and limited to a select chemical space. In this work, we explored if machine learning methods could improve the COSMO-SAC model and bridge density functional theory calculations to liquid phase thermodynamic properties. Our data-driven approach uses existing databases for sigma-profiles of organic solvents and reconciles their methodological differences via ensemble averaging. First, an optimal machine learning model is constructed for each dataset. Our machine learning algorithms use the sigma-profile as an input feature to predict binary mixtures' activity coefficients using multi-output regression. Each dataset uses different choices of functionals, methods, and basis sets. Therefore, our ensemble model attempts to predict corrected activity coefficients given the combination of all the model outputs. The activity coefficients used for training are generated using the COSMO-SAC model. This approach enables the extraction of meaningful information from the existing datasets to improve the COSMO-SAC model for obtaining thermodynamic properties of electrolyte mixtures. With the liquid phase activities, we can identify electrolyte mixtures that meet desired phase equilibria conditions.

Thermodynamics↗

High-Performance Semiempirical Excited-State Molecular Dynamics Powered by Graphics Processing Units

Here, this Letter introduces excited-state molecular dynamics in PYSEQM, a GPU-accelerated semiempirical quantum chemistry engine implemented in PyTorch. The new module enables Born–Oppenheimer molecular dynamics (BOMD) using configuration-interaction singles and random phase approximation for excited states, allowing long trajectories and large statistical ensembles to be simulated efficiently on a single GPU. We also implement an extended Lagrangian excited-state BOMD (XL-ESMD) scheme that propagates auxiliary electronic variables, enabling relaxed ground and excited-state convergence thresholds without compromising energy conservation. The excited-state BOMD implementation scales smoothly from small chromophores to a nearly 900-atom dendrimer (taking 6.5 s per MD step). PYSEQM also supports batched execution, allowing many geometries or trajectories to be evaluated in a single GPU launch, substantially increasing throughput and making ensemble-based protocols routine. As a demonstration, we compute absorption, emission, and infrared spectra from trajectories propagated on the ground and first excited states. The XL-ESMD scheme yields identical spectra at significantly lower computational cost, establishing the role of extended Lagrangian based dynamics for efficient excited-state BOMD simulations. Beyond raw performance, PYSEQM’s PyTorch foundation provides automatic differentiation for forces, efficient GPU batching, and seamless interfacing with machine learning models. These capabilities position PYSEQM as a practical platform for machine learning-augmented excited-state dynamics and lay the foundation for future data-driven nonadiabatic excited-state dynamics modeling of ultrafast spectroscopic probes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A unified ensemble soil moisture dataset across the continental United States

Abstract A unified ensemble soil moisture (SM) package has been developed over the Continental United States (CONUS). The data package includes 19 products from land surface models, remote sensing, reanalysis, and machine learning models. All datasets are unified to a 0.25-degree and monthly spatiotemporal resolution, providing a comprehensive view of surface SM dynamics. The statistical analysis of the datasets leverages the Koppen-Geiger Climate Classification to explore surface SM’s spatiotemporal variabilities. The extracted SM characteristics highlight distinct patterns, with the western CONUS showing larger coefficient of variation values and the eastern CONUS exhibiting higher SM values. Remote sensing datasets tend to be drier, while reanalysis products present wetter conditions. In-situ SM observations serve as the basis for wavelet power spectrum analyses to explain discrepancies in temporal scales across datasets facilitating daily SM records. This study provides a comprehensive soil moisture data package and an analysis framework that can be used for Earth system model evaluations and uncertainty quantification, quantifying drought impacts and land–atmosphere interactions and making recommendations for drought response planning.

54 ENVIRONMENTAL SCIENCES↗

Contrastive Machine Learning with Gamma Spectroscopy Data Augmentations for Detecting Shielded Radiological Material Transfers

Data analysis techniques can be powerful tools for rapidly analyzing data and extracting information that can be used in a latent space for categorizing observations between classes of data. Machine learning models that exploit learned data relationships can address a variety of nuclear nonproliferation challenges like the detection and tracking of shielded radiological material transfers. The high resource cost of manually labeling radiation spectra is a hindrance to the rapid analysis of data collected from persistent monitoring and to the adoption of supervised machine learning methods that require large volumes of curated training data. Instead, contrastive self-supervised learning on unlabeled spectra can enhance models that are built on limited labeled radiation datasets. This work demonstrates that contrastive machine learning is an effective technique for leveraging unlabeled data in detecting and characterizing nuclear material transfers demonstrated on radiation measurements collected at an Oak Ridge National Laboratory testbed, where sodium iodide detectors measure gamma radiation emitted by material transfers between the High Flux Isotope Reactor and the Radiochemical Engineering Development Center. Label-invariant data augmentations tailored for gamma radiation detection physics are used on unlabeled spectra to contrastively train an encoder, learning a complex, embedded state space with self-supervision. A linear classifier is then trained on a limited set of labeled data to distinguish transfer spectra between byproducts and tracked nuclear material using representations from the contrastively trained encoder. The optimized hyperparameter model achieves a balanced accuracy score of 80.30%. Any given model—that is, a trained encoder and classifier—shows preferential treatment for specific subclasses of transfer types. Regardless of the classifier complexity, a supervised classifier using contrastively trained representations achieves higher accuracy than using spectra when trained and tested on limited labeled data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automatic Loss Factor Modeling and Attribution on Unlabeled PV Energy Data

We present a novel approach for modeling the loss factors of photovoltaic power generation systems (PV systems). This method is a white-box machine learning model built on convex optimization that is fast, interpretable, and auditable. It takes as an input the measured daily energy produced by the system, over a multi-year period, and returns a multiplicative decomposition model of the daily energy signal and full attribution of the total energy loss to each feature. The methods section of this paper has two major components: (1) the description of the signal decomposition (SD) model, expressed in the SD framework, and (2) the attribution of total energy losses via Shapley values. We validate the method on synthetic and open-source data sets and compare to similar methods from the literature.

artificial intelligence↗

Comparison of Machine Learning-Based Predictive Models of the Nutrient Loads Delivered from the Mississippi/Atchafalaya River Basin to the Gulf of Mexico

Predicting nutrient loads is essential to understanding and managing one of the environmental issues faced by the northern Gulf of Mexico hypoxic zone, which poses a severe threat to the Gulf’s healthy ecosystem and economy. The development of hypoxia in the Gulf of Mexico is strongly associated with the eutrophication process initiated by excessive nutrient loads. Due to the complexities in the excessive nutrient loads to the Gulf of Mexico, it is challenging to understand and predict the underlying temporal variation of nutrient loads. The study was aimed at identifying an optimal predictive machine learning model to capture and predict nonlinear behavior of the nutrient loads delivered from the Mississippi/Atchafalaya River Basin (MARB) to the Gulf of Mexico. For this purpose, monthly nutrient loads (N and P) in tons were collected from US Geological Survey (USGS) monitoring station 07373420 from 1980 to 2020. Machine learning models—including autoregressive integrated moving average (ARIMA), gaussian process regression (GPR), single-layer multilayer perceptron (MLP), and a long short-term memory (LSTM) with the single hidden layer—were developed to predict the monthly nutrient loads, and model performances were evaluated by standard assessment metrics—Root Mean Square Error (RMSE) and Correlation Coefficient (R). The residuals of predictive models were examined by the Durbin–Watson statistic. The results showed that MLP and LSTM persistently achieved better accuracy in predicting monthly TN and TP loads compared to GPR and ARIMA. In addition, GPR models achieved slightly better test RMSE score than ARIMA models while their correlation coefficients are much lower than ARIMA models. Moreover, MLP performed slightly better than LSTM in predicting monthly TP loads while LSTM slightly outperformed for TN loads. Furthermore, it was found that the optimizer and number of inputs didn’t show effects on the LSTM performance while they exhibited impacts on MLP outcomes. This study explores the capability of machine learning models to accurately predict nonlinearly fluctuating nutrient loads delivered to the Gulf of Mexico. Further efforts focus on improving the accuracy of forecasting using hybrid models which combine several machine learning models with superior predictive performance for nutrient fluxes throughout the MARB.

54 ENVIRONMENTAL SCIENCES↗