Search NASA⌕ Search

SEARCH · Search NASA

Results for “empirical machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Comparative Analysis of Empirical and Machine Learning Models for Chla Extraction Using Sentinel-2 and Landsat OLI Data: Opportunities, Limitations, and Challenges

Remote retrieval of near-surface chlorophyll-a (Chla) concentration in small inland waters is challenging due to substantial optical interferences of various water constituents and uncertainties in the atmospheric correction (AC) process. Although various algorithms have been developed to estimate Chla from moderate-resolution terrestrial missions (∼10–60 m), the production of both accurate distribution maps and time series of Chla has proven challenging, limiting the use of remote analyses for lake monitoring. Here, we develop a support vector regression (SVR) model, which uses satellite-derived remote-sensing reflectance spectra () from Sentinel-2 and Landsat-8 images as input for Chla retrieval in a representative eutrophic prairie lake, Buffalo Pound Lake (BPL), Saskatchewan, Canada. Validated against in situ Chla from seven ice-free seasons (N ∼ 200; 2014–2020), the SVR model outperformed both locally tuned, -fed empirical models (Normalized Difference Chlorophyll Index, 2- and 3-band, and OC3) and Mixture Density Networks (MDNs) by 15–65%, while exhibiting comparable performance to a locally trained MDN, with an error of ∼35%. Comparison of Chla retrieval models, AC processors (iCOR, ACOLITE), and radiometric products (Rayleigh-corrected, surface, and top-of-atmosphere reflectance) showed that the best Chla maps and optimal time series (up to 100 mg m−3) were produced using a coupled SVR-iCOR system.

algal blooms↗

Machine Learning Emulators and Empirical Models Combining Climate and Global Crop Models for Seasonal Agricultural Production

We present results from several connected efforts to apply machine learning methods to estimates of seasonal agricultural production anomalies around the world. First, we apply the XGBoost Random Forest method to fit emulators that mimic global crop models participating in the Agricultural Model Intercomparison and Improvement Project (AgMIP) Global Gridded Crop Model Intercomparison (GGCMI). These are the same models used in the agricultural sector simulations of the Inter-Sectoral Impacts Model Intercomparison Project (ISIMIP). These emulators use 8 climate variables split across 5 sub-seasonal representations of the growing season for each ½ degree grid cell around the world for maize, wheat, rice and soybeans. Emulators are useful for estimating conditions that have not already been simulated by GGCMI (e.g., in a seasonal prediction model) and also to diagnose model differences and capabilities. For example, emulators of the pDSSAT maize model tend to be more reliant on mean temperatures than the LPJmL model, and few models have strong responses to cold extremes. Second, we use a similar XGBoost approach to fit empirical models for national production data for the top 20 producing countries according to the United Nations Food and Agricultural Organization (FAO). Models utilize both climate observations and the GGCM models as predictors, resulting in skillful models for many (but not all) top producing-countries. The patterns of climate and crop model features selected indicate regions and systems that are better or worse simulated by the GGCMs. For example, information in cold extreme predictors is often combined with GGCM output predictors to provide sensitivity that models may underrepresent.

machine learning↗

Remote Sensing of CDOM, CDOM Spectral Slope, and Dissolved Organic Carbon in the Global Ocean

A Global Ocean Carbon Algorithm Database (GOCAD) has been developed from over 500 oceanographic field campaigns conducted worldwide over the past 30 years including in situ reflectances and coincident satellite imagery, multi- and hyperspectral Chromophoric Dissolved Organic Matter (CDOM) absorption coefficients from 245–715 nm, CDOM spectral slopes in eight visible and ultraviolet wavebands, dissolved and particulate organic carbon (DOC and POC, respectively), and inherent optical, physical, and biogeochemical properties. From field optical and radiometric data and satellite measurements, several semi-analytical, empirical, and machine learning algorithms for retrieving global DOC, CDOM, and CDOM slope were developed, optimized for global retrieval, and validated. Global climatologies of satellite-retrieved CDOM absorption coefficient and spectral slope based on the most robust of these algorithms lag seasonal patterns of phytoplankton biomass belying Case 1 assumptions, and track terrestrial runoff on ocean basin scales. Variability in satellite retrievals of CDOM absorption and spectral slope anomalies are tightly coupled to changes in atmospheric and oceanographic conditions associated with El Niño Southern Oscillation (ENSO), strongly covary with the multivariate ENSO index in a large region of the tropical Pacific, and provide insights into the potential evolution and feedbacks related to sea surface dissolved carbon in a warming climate. Further validation of the DOC algorithm developed here is warranted to better characterize its limitations, particularly in mid-ocean gyres and the southern oceans.

Dissolved organic carbon↗

Robust Algorithm for Estimating Total Suspended Solids (TSS) in Inland and Nearshore Coastal Waters

One of the challenging tasks in modern aquatic remote sensing is the retrieval of near-surface concentrations of Total Suspended Solids (TSS). This study aims to present a Statistical, inherent Optical property (IOP) -based, and muLti-conditional Inversion proceDure (SOLID) for enhanced retrievals of satellite-derived TSS under a wide range of in-water bio-optical conditions in rivers, lakes, estuaries, and coastal waters. In this study, using a large in situ database (N > 3500), the SOLID model is devised using a three-step procedure: (a) water-type classification of the input remote sensing reflectance (R(sub rs)), (b) retrieval of particulate backscattering (b(sub bp)) in the red or near-infrared (NIR) regions using semi-analytical, machine-learning, and empirical models, and (c) estimation of TSS from b(sub bp) via water-type-specific empirical models. Using an independent subset of our in situ data (N = 2729) with TSS ranging from 0.1 to 2626.8 [g/m (exp 3)], the SOLID model is thoroughly examined and compared against several state-of-the-art algorithms (Miller and McKee, 2004; Nechad et al., 2010; Novoa et al., 2017; Ondrusek et al., 2012; Petus et al., 2010). We show that SOLID outperforms all the other models to varying degrees, i.e., from 10 to > 100%, depending on the statistical attributes (e.g., global versus water-type-specific metrics). For demonstration purposes, the model is implemented for images acquired by the MultiSpectral Imager aboard Sentinel-2A/B over the Chesapeake Bay, San-Francisco-Bay-Delta Estuary, Lake Okeechobee, and Lake Taihu. To enable generating consistent, multimission TSS products, its performance is further extended to, and evaluated for, other missions, such as the Ocean and Land Color Instrument (OLCI), Moderate Resolution Imaging Spectroradiometer (MODIS), Visible Infrared Imaging Radiometer Suite (VIIRS), and Operational Land Imager (OLI). Sensitivity analyses on uncertainties induced by the atmospheric correction indicate that 10% uncertainty in Rrs leads to < 20% uncertainty in TSS retrievals from SOLID. While this study suggests that SOLID has a potential for producing TSS products in global coastal and inland waters, our statistical analysis certainly verifies that there is still a need for improving retrievals across a wide spectrum of particle loads.

Total suspended solids↗

Myths and legends in learning classification rules

A discussion is presented of machine learning theory on empirically learning classification rules. Six myths are proposed in the machine learning community that address issues of bias, learning as search, computational learning theory, Occam's razor, universal learning algorithms, and interactive learning. Some of the problems raised are also addressed from a Bayesian perspective. Questions are suggested that machine learning researchers should be addressing both theoretically and experimentally.

Buntine, Wray↗

Myths and legends in learning classification rules

This paper is a discussion of machine learning theory on empirically learning classification rules. The paper proposes six myths in the machine learning community that address issues of bias, learning as search, computational learning theory, Occam's razor, 'universal' learning algorithms, and interactive learnings. Some of the problems raised are also addressed from a Bayesian perspective. The paper concludes by suggesting questions that machine learning researchers should be addressing both theoretically and experimentally.

Buntine, Wray↗

A Machine Learning Approach to Jet-Surface Interaction Noise Modeling

This paper investigates using machine learning to rapidly develop empirical models suitable for system-level aircraft noise studies. In particular, machine learning is used to train a neural network to predict the noise spectra produced by a round jet near a surface over a range of surface lengths, surface standoff distances, jet Mach numbers, and observer angles. These spectra include two sources, jet-mixing noise and jet-surface interaction (JSI) noise, with different scale factors as well as surface shielding and reflection effects to create a multi- dimensional problem. A second model is then trained using data from three rectangular nozzles to include nozzle aspect ratio in the spectral prediction. The training and validation data are from an extensive jet-surface interaction noise database acquired at the NASA Glenn Research Center's Aero-Acoustic Propulsion Laboratory. Although the number of training and validation points is small compared a typical machine learning application, the results of this investigation show that this approach is viable if the underlying data are well behaved.

Brown, Cliff↗

Review of Solar Energetic Particle Models

Solar Energetic Particle (SEP) events are interesting from a scientific perspective as they are the product of a broad set of physical processes from the corona out through the extent of the heliosphere, and provide insight into processes of particle acceleration and transport that are widely applicable in astrophysics. From the operations perspective, SEP events pose a radiation hazard for aviation, electronics in space, and human space exploration, in particular for missions outside of the Earth’s protective magnetosphere including to the Moon and Mars. Thus, it is critical to improve the scientific understanding of SEP events and use this understanding to develop and improve SEP forecasting capabilities to support operations. Many SEP models exist or are in development using a wide variety of approaches and with differing goals. These include computationally intensive physics-based models, fast and light empirical models, machine learning-based models, and mixed-model approaches. The aim of this paper is to summarize all of the SEP models currently developed in the scientific community, including a description of model approach, inputs and outputs, free parameters, and any published validations or comparisons with data.

Kathryn Whitman↗

Machine Learning for the Validation of Expert-Elicited Causal Risk Diagrams

Exposure to spaceflight poses risk to human health in complex ways. To help manage this risk, the Human Systems Risk Board (HSRB) at the National Aeronautics and Space Administration (NASA) maintains a set of causal diagrams that attempt to explain how spaceflight hazards generate health risks and lead to adverse outcomes both in-mission, immediately post-mission, and over the long term. These causal risk diagrams are formulated as directed acyclic graphs (DAGs) and can function as knowledge graphs of connected risks and outcomes. These DAGs have proven useful for communication, and, through network analysis, have allowed for the identification of structurally important factors in the risk network. However, the utility these DAGs provide is directly proportional to their verisimilitude, making assessment of this trait using empirical data – whether from actual human spaceflight or various spaceflight analogue exposures and model organisms – a high priority. In this research we explore the use of machine learning algorithms to learn DAG structure from empirical data as a means of evaluating human-elicited DAG structures. To do so, we test several different graph structure-learning algorithms on data concerning changes in the bones of rats and mice after exposure to either spaceflight or a spaceflight analogue. We explore potential methods for indexing the similarity between each algorithm’s output DAG with all the others and with that of the expert-elicited DAG. We discuss next steps in this ongoing line of research and open science initiatives underway to complete them.

directed acyclic graphs↗

Measuring Constraint-Set Utility for Partitional Clustering Algorithms

Clustering with constraints is an active area of machine learning and data mining research. Previous empirical work has convincingly shown that adding constraints to clustering improves the performance of a variety of algorithms. However, in most of these experiments, results are averaged over different randomly chosen constraint sets from a given set of labels, thereby masking interesting properties of individual sets. We demonstrate that constraint sets vary significantly in how useful they are for constrained clustering; some constraint sets can actually decrease algorithm performance. We create two quantitative measures, informativeness and coherence, that can be used to identify useful constraint sets. We show that these measures can also help explain differences in performance for four particular constrained clustering algorithms.

constraints↗

A Machine Learning Approach to Predict Martensitic Transition Temperatures for Shape Memory Alloys

Shape memory alloys (SMAs) are a unique class of materials with several remarkable properties including shape recovery, superelasticity, etc. Especially important for many NASA applications is the ability to tune the martensitic phase transition temperature by varying the alloy composition. Nickel-titanium (NiTi) based alloys are the most widely studied of this class, with compositions involving ternary, quaternary, or higher additions being considered. Over the past several years, a significant database of SMA properties has been assembled by NASA researchers. Such a database is ideal for data science-based approaches including machine learning. We present results from a developed machine learning model capable of accurately predicting the transition temperature of SMAs across a wide range of compositions. Our model has the added benefit of interpretability and even provides confidence intervals for our predictions. This model will make rapid screening and design of new SMA materials possible. Predictions from the machine learning model can be validated by empirical and/or atomistic scale modeling.

Shreyas Honrao↗

Paradigms for machine learning

Five paradigms are described for machine learning: connectionist (neural network) methods, genetic algorithms and classifier systems, empirical methods for inducing rules and decision trees, analytic learning methods, and case-based approaches. Some dimensions are considered along with these paradigms vary in their approach to learning, and the basic methods are reviewed that are used within each framework, together with open research issues. It is argued that the similarities among the paradigms are more important than their differences, and that future work should attempt to bridge the existing boundaries. Finally, some recent developments in the field of machine learning are discussed, and their impact on both research and applications is examined.

Schlimmer, Jeffrey C.↗

Celebrating 10 Years of the Sub-Seasonal to Seasonal Prediction Project and Looking to the Future

The conference clearly demonstrated the increasing interest and growth of the scientific community working on the development and application of sub-seasonal to seasonal prediction since the start of the World Weather Research Programme (WWRP)/World Climate Research Programme (WCRP) sub-seasonal to seasonal (S2S) prediction project in 2013. The conference, which was held at the University of Reading (United Kingdom), was organized into three main themes as briefly summarized below, with eleven invited talks, 74 oral contributed talks, and 101 posters. The conference also included a two-hour breakout session, wherein eight groups discussed the current state and prospect for S2S prediction, and an early career researcher event. A summary of these discussions and recommendations is presented below. The conference web page (https://research.reading.ac.uk/s2s-summit2023/) is archived at the University of Reading. Introductory comments by representatives of the World Meteorological Organization (WMO) WWRP and WCRP emphasized the importance of the weather–climate linkage, targeted by S2S forecasts (from 2 weeks to a season ahead), addressing the challenges of creating “end-to-end” forecasts that encompass the entire climate-services chain from the prediction science and forecast, to the development and issuing of forecast products tailored to informing user-decisions. They also emphasized the efficacy of multi-model ensemble efforts and databases to foster collaborations internationally and between operational centres and academia. Although the WWRP/WCRP S2S project comes to an end in 2023, S2S prediction will remain an important focus for WWRP and WCRP. In WWRP, a new project called SAGE (Sub-seasonal to seasonal predictions for Agriculture and Environment) will start in 2024. Another important legacy of the S2S project will be the maintenance of the S2S database (Vitart et al. 2017) and the establishment of a WMO Lead Center for sub-seasonal prediction multi-model ensemble (LC-SSPMME) which will provide real-time multi-model S2S climate information. In two keynote presentations, Prof. Brian Hoskins (University of Reading) and Dr. Gilbert Brunet (Australian Bureau of Meteorology) discussed the potential of S2S predictability and the ongoing journey for understanding and improving these predictions. This conference was a sequel to the International Conference on Sub-seasonal to Seasonal Prediction (Robertson et al., 2014) which took place in College Park (Maryland, USA) in February 2014 to celebrate the start of the WWRP/WCRP S2S project, and to WCRP and WWRP conferences in Boulder, USA, in 2018 (Merryfield et al., 2020). A significant development compared to the previous S2S conferences was the large number of presentations on research to operation (R2O) and S2S applications and on the use of artificial intelligence and machine learning (AI/ML) methods for S2S prediction. Some of these methods provide empirical S2S forecasts which are competitive with state-of-the-art dynamical models. Other presentations demonstrated that AI/ML can provide alternative calibration of dynamical model outputs to traditional methods. Several talks and posters highlighted the increasing use of AI/ML, including deep learning, in S2S forecast post-processing and using AI to identify higher flow-dependent skill. Finally, some presentations demonstrated the value of AI/ML methods for a better understanding of S2S sources of predictability and attribution of extreme events.

S. J. Woolnough↗

A Comprehensive Machine Learning Study to Classify Precipitation Type over Land from Global Precipitation Measurement Microwave Imager (GPM-GMI) Measurements

Precipitation type is a key parameter used for better retrieval of precipitation characteristics as well as to understand the cloud–convection–precipitation coupling processes. Ice crystals and water droplets inherently exhibit different characteristics in different precipitation regimes (e.g., convection, stratiform), which reflect on satellite remote sensing measurements that help us distinguish them. The Global Precipitation Measurement (GPM) Core Observatory’s microwave imager (GMI) and dual-frequency precipitation radar (DPR) together provide ample information on global precipitation characteristics. As an active sensor, the DPR provides an accurate precipitation type assignment, while passive sensors such as the GMI are traditionally only used for empirical understanding of precipitation regimes. Using collocated precipitation type flags from the DPR as the “truth”, this paper employs machine learning (ML) models to train and test the predictability and accuracy of using passive GMI-only observations together with ancillary information from a reanalysis and GMI surface emissivity retrieval products. Out of six ML models, four simple ones (support vector machine, neural network, random forest, and gradient boosting) and the 1-D convolutional neural network (CNN) model are identified to produce 90–94% prediction accuracy globally for five types of precipitation (convective, stratiform, mixture, no precipitation, and other precipitation), which is much more robust than previous similar effort. One novelty of this work is to introduce data augmentation (subsampling and bootstrapping) to handle extremely unbalanced samples in each category. A careful evaluation of the impact matrices demonstrates that the polarization difference (PD), brightness temperature (Tc) and surface emissivity at high-frequency channels dominate the decision process, which is consistent with the physical understanding of polarized microwave radiative transfer over different surface types, as well as in snow and liquid clouds with different microphysical properties. Furthermore, the view-angle dependency artifact that the DPR’s precipitation flag bears with does not propagate into the conical-viewing GMI retrievals. This work provides a new and promising way for future physics-based ML retrieval algorithm development.

machine learning/artificial intelligence↗

Machine Learning the COSMO Model for Predicting Thermodynamics of Electrolyte Mixtures

Bottom-up design of electrolyte mixtures for battery systems requires predicting macro thermodynamic properties from molecular constituents. For instance, molten salt electrolyte batteries require conditions far above room temperature to operate. Therefore, discovering mixtures with increasingly lower eutectic melting points is desirable. A model that can approximate chemical activity is a valuable tool to search through the vast compositional design space. Machine learning can predict properties of materials such as vibrational free energies, electronic energy gaps, and thermal conductivities. Moreover, they can learn physical models such as interatomic potentials. The COSMO-SAC model uses theory and empirical parameterization to predict liquid-vapor and liquid-solid properties using first-principles calculations. However, obtaining activity coefficients required for parameterizing the COSMO-SAC model is costly and limited to a select chemical space. In this work, we explored if machine learning methods could improve the COSMO-SAC model and bridge density functional theory calculations to liquid phase thermodynamic properties. Our data-driven approach uses existing databases for sigma-profiles of organic solvents and reconciles their methodological differences via ensemble averaging. First, an optimal machine learning model is constructed for each dataset. Our machine learning algorithms use the sigma-profile as an input feature to predict binary mixtures' activity coefficients using multi-output regression. Each dataset uses different choices of functionals, methods, and basis sets. Therefore, our ensemble model attempts to predict corrected activity coefficients given the combination of all the model outputs. The activity coefficients used for training are generated using the COSMO-SAC model. This approach enables the extraction of meaningful information from the existing datasets to improve the COSMO-SAC model for obtaining thermodynamic properties of electrolyte mixtures. With the liquid phase activities, we can identify electrolyte mixtures that meet desired phase equilibria conditions.

Thermodynamics↗

Hierarchical screening for Li-based solid electrolytes using fast, interpretable machine-learned potentials

Li-based solid-state electrolyte materials enable safer, all-solid-state batteries but the computational search for candidates with favorable stability and Li-ion conductivity is challenging due to the size of the search space and the cost of evaluating transport properties with ab initio methods. The prohibitive cost of high-throughput screening with DFT has lead to the development of surrogate models using geometric analysis, empirical potentials, and descriptors for ionic transport. Here, I will discuss a hierarchical screening approach for identifying promising materials using a combination of density functional theory, bond-valence methods, and machine learning potentials generated with the Ultra-Fast Force Fields (UF3) framework. We show how the inexpensive bond-valence method can be used to guide the generation of training samples for machine learning, in addition to filtering candidates. Finally, we apply the hierarchical workflow to screen for ionic conductivity across a database of Li-containing compounds.

Materials discovery↗

Hierarchical Screening for Li-Based Solid Electrolytes Using Fast, Interpretable Machine-Learned Potentials

Li-based solid-state electrolyte materials enable safer, all-solid-state batteries but the computational search for candidates with favorable stability and Li-ion conductivity is challenging due to the size of the search space and the cost of evaluating transport properties with ab initio methods. The prohibitive cost of high-throughput screening with DFT has lead to the development of surrogate models using geometric analysis, empirical potentials, and descriptors for ionic transport. Here, I will discuss a hierarchical screening approach for identifying promising materials using a combination of density functional theory, bond-valence methods, and machine learning potentials generated with the Ultra-Fast Force Fields (UF3) framework. We show how the inexpensive bond-valence method can be used to guide the generation of training samples for machine learning, in addition to filtering candidates.

Materials discovery↗