Search NASA⌕ Search

SEARCH · Search NASA

Results for “database for machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Mapping wall-to-wall fractional cover of Arctic tundra plant functional types in Alaska using 20-m spatial resolution satellite imagery and harmonized plot observations

Estimates of fractional cover (fCover) across given land surfaces are used to assess, and often model, vegetation composition and diversity, which are crucial for understanding the health and functioning of terrestrial ecosystems. Remote sensing provides a useful means for scaling local, plot-measured fCover estimates to regional scales. Leveraging a recently synthesized and harmonized plot database, this study generated wall-to-wall maps of fCover for six Alaskan-Arctic plant functional types (PFT), including non-vascular plants, forbs, graminoids, and deciduous and evergreen shrubs, using 20-m satellite data (Sentinel-1, Sentinel-2, ArcticDEM) using a machine learning regression approach, specifically the random forest (RF) algorithm, which is well-suited for handling nonlinear relationships and high-dimensional satellite datasets. This study additionally addressed the spatio-temporal inconsistencies e.g., sampling scale, plot size, and collection year in plot measured fCover by adopting a multivariate outlier detection approach—Cook’s distance—to identify high-quality plots for model training and validation. Our approach achieves high accuracy (R 2 = 0.59–0.93, root mean squared errors = 0.02–0.10 for all PFTs) between plot-observed and satellite-derived fCover when using high-quality plot samples. The mapped fCover characterizes the spatial patterns of different PFTs across the tundra biome at a 20-m resolution, providing key information needed for improved representation of Arctic tundra vegetation in terrestrial biosphere models to better understand climate-vegetation feedback across the Arctic tundra.

Arctic tundra↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Automatic cataloguing and characterization of Earth science data using SE-trees

In the future, NASA's Earth Observing System (EOS) platforms will produce enormous amounts of remote sensing image data that will be stored in the EOS Data Information System. For the past several years, the Intelligent Data Management group at Goddard's Information Science and Technology Office has been researching techniques for automatically cataloguing and characterizing image data (ADCC) from EOS into a distributed database. At the core of the approach, scientists will be able to retrieve data based upon the contents of the imagery. The ability to automatically classify imagery is key to the success of contents-based search. We report results from experiments applying a novel machine learning framework, based on Set-Enumeration (SE) trees, to the ADCC domain. We experiment with two images: one taken from the Blackhills region in South Dakota; and the other from the Washington DC area. In a classical machine learning experimentation approach, an image's pixels are randomly partitioned into training (i.e. including ground truth or survey data) and testing sets. The prediction model is built using the pixels in the training set, and its performance is estimated using the testing set. With the first Blackhills image, we perform various experiments achieving an accuracy level of 83.2 percent, compared to 72.7 percent using a Back Propagation Neural Network (BPNN) and 65.3 percent using a Gaussain Maximum Likelihood Classifier (GMLC). However, with the Washington DC image, we were only able to achieve 71.4 percent, compared with 67.7 percent reported for the BPNN model and 62.3 percent for the GMLC.

Rymon, Ron↗

Virtual Assistant for First Responders Using Natural Language Understanding and Optical Character Recognition

Commercial deep learning capabilities are available for many applications such as computer vision processing and intelligent chat bots. The Google Cloud Platform product Google Dialogflow provides lifelike conversational artificial intelligence (AI) using machine learning (ML) to generate natural conversations between computers and humans. This ML utilizes natural language understanding (NLU) to recognize a user’s intent and extracts key information into a form of entities. We have developed a user-friendly application through understanding the hazardous material database, first aid safety guidelines and observing the process of first responders who access this information in the field. We created the Trusted and Explainable Artificial Intelligence for Saving Lives (TruePAL) virtual assistant using Dialogflow1 and TensorFlow2 paired with EasyOCR.3 The chatbot supports first responders by providing voice interaction which helps limit additional steps such as browsing through multiple categories when searching for information. Using feedback from our field interviews, the voice interface has been developed to enable the first responder to focus on the immediate emergency. With less distractions, the first responder is able to engage the incident more effectively. The partial hands-free TruePAL chatbot assistant improves the accessibility to the correct guidance by an average of 1.9 seconds compared to the widely used application, NIH WISER, which requires full attention to operate. We combined this intelligent chatbot with a separate visual processing capability to produce hazardous signage analysis and generate the proper guidance for first responders. With the evolving functionality of AI tools, the use of virtual assistants in first responder technology will be an advancement, benefiting the safety of both first responders and civilians.

Chow, Edward↗

Autonomous Synthesis and Inverse Design of Electrochromic Polymers with High Efficiency and Accuracy

Here, the design and synthesis of functional polymers, aimed at targeted properties through specific structures, have long been challenged by their complex and often nonlinear structure–property relationships. Key processes, including knowledge accumulation for predictive design and experimental refinement and validation, are traditionally labor-insensitive and time-consuming, making it difficult to balance accuracy and efficiency. Here, we introduce an accelerated, autonomous system for the on-demand synthesis of electronic polymers that achieves the desired electrochromic functionality with high accuracy and efficiency. Our approach leverages large language model-assisted data mining, a physics-informed copolymer machine learning model, and an AI-driven autonomous robotic workflow in the Polybot lab. Within 72 h, Polybot autonomously synthesized electrochromic polymers (ECPs) with targeted, previously-unreported color values, including green polymers with specific absorption profiles, precisely fine-tuning copolymer structures with a 5% step size in comonomer composition within a three-monomer system. A publicly accessible ECP informatics database has also been created to foster knowledge exchange.

AI-driven Robotic Lab↗

Data Science-Driven Discovery of Multimetallic Oxygen-cycle Electrocatalysts for Enhanced Energy Conversion

The overarching objective of this effort has been to combine state-of-the-art data science techniques, first principles analyses, and molecular-level characterization of electrocatalyst structure and reactivity to identify both in-situ mechanisms for degradation and transformation of electrocatalysts with highly complex catalytic structures and the impact of these transformations on catalytic activity. The primary catalysts of interest have been multielemental alloys, including high entropy alloys (HEA’s), which are characterized by a high degree of disorder and up to 20 different elements within a single nanoparticle. We have applied these strategies primarily to energy-critical oxygen cycle electrocatalytic reactions, including oxygen reduction (ORR), but we have also considered extensions to non-electrochemical chemistries such as ammonia synthesis and decomposition. We have made strong progress in the development of computational methods on both the level of machine learning methods development as well as first principles-based treatments of HEA’s, and we have leveraged these insights to propose promising HEA catalysts for the ORR. On the experimental side, we developed new HEA synthesis and characterization protocols relevant to these reactions and developed a database combining our experimental results with corresponding computational tools.

36 MATERIALS SCIENCE↗

A Chlorophyll-a Algorithm for Landsat-8 Based on Mixture Density Networks

Retrieval of aquatic biogeochemical variables, such as the near-surface concentration of chlorophyll-a (Chla) in inland and coastal waters via remote observations, has long been regarded as a challenging task. This manuscript applies Mixture Density Networks (MDN) that use the visible spectral bands available by the Operational Land Imager (OLI) aboard Landsat-8 to estimate Chla. We utilize a database of co-located in situ radiometric and Chla measurements (N = 4,354), referred to as Type A data, to train and test an MDN model (MDN(A)). This algorithm’s performance, having been proven for other satellite missions, is further evaluated against other widely used machine learning models (e.g., support vector machines), as well as other domain-specific solutions (OC3), and shown to offer significant advancements in the field. Our performance assessment using a held-out test data set suggests that a 49% (median) accuracy with near-zero bias can be achieved via the MDN(A) model, offering improvements of 20 to 100% in retrievals with respect to other models. The sensitivity of the MDN(A) model and benchmarking methods to uncertainties from atmospheric correction (AC) methods, is further quantified through a semi-global matchup dataset (N = 3,337), referred to as Type B data. To tackle the increased uncertainties, alternative MDN models (MDN(B)) are developed through various features of the Type B data (e.g., Rayleigh-corrected reflectance spectra ρ(s)). Using held-out data, along with spatial and temporal analyses, we demonstrate that these alternative models show promise in enhancing the retrieval accuracy adversely influenced by the AC process. Results lend support for the adoption of MDN(B) models for regional and potentially global processing of OLI imagery, until a more robust AC method is developed. Index Terms—Chlorophyll-a, coastal water, inland water, Landsat-8, machine learning, ocean color, aquatic remote sensing.

Brandon Smith↗

Simultaneous Retrieval of Selected Optical Water Quality Indicators From Landsat-8, Sentinel-2, and Sentinel-3

Constructing multi-source satellite-derived water quality (WQ) products in inland and nearshore coastal waters from the past, present, and future missions is a long-standing challenge. Despite inherent differences in sensors’ spectral capability, spatial sampling, and radiometric performance, research efforts focused on formulating, implementing, and validating universal WQ algorithms continue to evolve. This research extends a recently developed machine-learning (ML) model, i.e., Mixture Density Networks (MDNs) (Pahlevan et al., 2020; Smith et al., 2021), to the inverse problem of simultaneously retrieving WQ indicators, including chlorophyll-a (Chla), Total Suspended Solids (TSS), and the absorption by Colored Dissolved Organic Matter at 440 nm (a cdom (440)), across a wide array of aquatic ecosystems. We use a database of in situ measurements to train and optimize MDN models developed for the relevant spectral measurements (400–800 nm) of the Operational Land Imager (OLI), MultiSpectral Instrument (MSI), and Ocean and Land Color Instrument (OLCI) aboard the Landsat-8, Sentinel-2, and Sentinel-3 missions, respectively. Our two performance assessment approaches, namely hold-out and leave-one-out, suggest significant, albeit varying degrees of improvements with respect to second-best algorithms, depending on the sensor and WQ indicator (e.g., 68%, 75%, 117% improvements based on the hold-out method for Chla, TSS, and a cdom (440), respectively from MSI-like spectra). Using these two assessment methods, we provide theoretical upper and lower bounds on model performance when evaluating similar and/or out-of-sample datasets. To evaluate multi-mission product consistency across broad spatial scales, map products are demonstrated for three near-concurrent OLI, MSI, and OLCI acquisitions. Overall, estimated TSS and a cdom (440) from these three missions are consistent within the uncertainty of the model, but Chla maps from MSI and OLCI achieve greater accuracy than those from OLI. By applying two different atmospheric correction processors to OLI and MSI images, we also conduct matchup analyses to quantify the sensitivity of the MDN model and best-practice algorithms to uncertainties in reflectance products. Our model is less or equally sensitive to these uncertainties compared to other algorithms. Recognizing their uncertainties, MDN models can be applied as a global algorithm to enable harmonized retrievals of Chla, TSS, and a cdom (440) in various aquatic ecosystems from multi-source satellite imagery. Local and/or regional ML models tuned with an apt data distribution (e.g., a subset of our dataset) should nevertheless be expected to outperform our global model.

Machine learning↗

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),↗

Managing the Digital Thread for Structural Applications With Fit for Purpose Materials

With the increased emphasis on reducing the cost and time to market of new materials, the need for analytical tools that enable the virtual design and optimization of materials throughout their processing - internal structure - property - performance envelope, along with the capturing and storing of the associated material and model information across its lifecycle, has become critical. This need is also fueled by the demands for higher efficiency in material testing; consistency, quality and traceability of data; product design; engineering analysis; as well as control of access to proprietary or sensitive information. Consequently, at NASA Glenn Research Center a robust information management system that manages the digital thread across the full material life (i.e., capture, analysis, maintenance, and dissemination of data) cycle directed at the design of ‘fit-for-purpose materials’ is under development. To this end the Application Table has been incorporated within NASA Glenn Research Center’s ICME Information Management framework within the ANSYS Granta MI tool. The Application Table provides a place where material and structural application information/requirements can be linked to marry the “design-the-material” (structural engineering) and the “design-with-material” (material science) paradigms and thereby enable application-driven design and optimization of materials and structures. In additional several associated toolsets, specifically: AIMAOS (Automated Information Management Across Organizations and Scales), Py MILab, and JARIMIS (Just A Rather Intelligent Material Interrogation System) are also under development to assist in the judicious automation of this process. AIMOAS offers users an interactive graphical user interface for connecting material information management systems with both commercial and in-house simulation tools at various length scales to enable such automation in the handoff across scales and maintenance of material digital twins and the digital thread. Py MILab, is an automatic framework for the capture, analysis, maintenance, and storage of material test data. Py MILab uses a modular approach for capturing raw data, analyzing the data, and storing the data in a database, interfaced by neutral file structures, to promote plug-and-play capabilities for various analysis types. Finally, JARIMIS is an expert system that integrates various materials informatics tools (e.g., MicroNet, Surrogate ML models, ANSYS Granta MI, etc.) to enable inverse design of materials and facilitate the application of machine learning (ML) and data science with human in the loop decision making to rapidly discover and optimize new materials.

Digital Transformation↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

Robust Algorithm for Estimating Total Suspended Solids (TSS) in Inland and Nearshore Coastal Waters

One of the challenging tasks in modern aquatic remote sensing is the retrieval of near-surface concentrations of Total Suspended Solids (TSS). This study aims to present a Statistical, inherent Optical property (IOP) -based, and muLti-conditional Inversion proceDure (SOLID) for enhanced retrievals of satellite-derived TSS under a wide range of in-water bio-optical conditions in rivers, lakes, estuaries, and coastal waters. In this study, using a large in situ database (N > 3500), the SOLID model is devised using a three-step procedure: (a) water-type classification of the input remote sensing reflectance (R(sub rs)), (b) retrieval of particulate backscattering (b(sub bp)) in the red or near-infrared (NIR) regions using semi-analytical, machine-learning, and empirical models, and (c) estimation of TSS from b(sub bp) via water-type-specific empirical models. Using an independent subset of our in situ data (N = 2729) with TSS ranging from 0.1 to 2626.8 [g/m (exp 3)], the SOLID model is thoroughly examined and compared against several state-of-the-art algorithms (Miller and McKee, 2004; Nechad et al., 2010; Novoa et al., 2017; Ondrusek et al., 2012; Petus et al., 2010). We show that SOLID outperforms all the other models to varying degrees, i.e., from 10 to > 100%, depending on the statistical attributes (e.g., global versus water-type-specific metrics). For demonstration purposes, the model is implemented for images acquired by the MultiSpectral Imager aboard Sentinel-2A/B over the Chesapeake Bay, San-Francisco-Bay-Delta Estuary, Lake Okeechobee, and Lake Taihu. To enable generating consistent, multimission TSS products, its performance is further extended to, and evaluated for, other missions, such as the Ocean and Land Color Instrument (OLCI), Moderate Resolution Imaging Spectroradiometer (MODIS), Visible Infrared Imaging Radiometer Suite (VIIRS), and Operational Land Imager (OLI). Sensitivity analyses on uncertainties induced by the atmospheric correction indicate that 10% uncertainty in Rrs leads to < 20% uncertainty in TSS retrievals from SOLID. While this study suggests that SOLID has a potential for producing TSS products in global coastal and inland waters, our statistical analysis certainly verifies that there is still a need for improving retrievals across a wide spectrum of particle loads.

Total suspended solids↗

Predictive Chemical Kinetic Modeling: Where We Succeed, Where We Struggle, and What Comes Next

Chemical kinetic modeling plays a foundational role in fields ranging from energy to environmental science, pharmaceuticals, and advanced materials. The past two decades have seen remarkable progress, particularly in modeling gas-phase reactions for thermochemical processes, leading to impactful industrial applications such as steam cracking and air quality management. However, new challenges are emerging. The successful development of systematic methodologies for the description of gas-phase kinetics opens the possibility to apply the same approach to the study of more challenging systems. Here, we review recent advances, including ab initio transition state theory-based master equation estimation of elementary rates, automated mechanism generation, machine-learning-assisted kinetics, and uncertainty quantification, and discuss the advances needed to apply the same methodological approach in areas such as heterogeneous catalysis, electrochemistry, liquid-phase and solid-state reactivity, and multiscale model integration. We advocate for the development of targeted tools, especially methods that go beyond empirical tuning toward first-principles-based predictions. We highlight the need for accessible software and AIaugmented workflows to democratize modeling for industry and academia alike. In this perspective, we call attention to not only what has worked but also what remains unsolved, advocating to avoid overemphasizing successes in scientific works at the expense of realism. The next decade should focus on predictive capability, physical accuracy, and community infrastructure (e.g., databases and services) to enable innovation across diverse fields. We argue that kinetic modeling, properly equipped, can accelerate discovery far beyond its traditional domains.

ab initio calculations↗

Machine learning for the redox potential prediction of molecules in organic redox flow battery

Here, organic redox flow batteries (ORFB) are recognized as an innovative technology for the large-scale storage of renewable energy. The redox potential of organic redox-active molecules plays a vital role in their performance. Advanced screening techniques like high-throughput experiment and machine learning (ML) have significantly enhanced organic material performance and transformed the field of ORFB. However, the scarcity of experimental data poses a considerable challenge for ML model development in this domain. In our study, we developed lightweight graph-based Gaussian process regression (GPR) models with GPU-accelerated marginalized graph kernel and hybrid kernel to predict the redox potentials of organic redox-active molecules for ORFBs, specifically focusing on small datasets. To evaluate model accuracy, we created a new experimental database of organic redox-active molecules by the data from hundreds of published papers and assembled previous computational datasets. We also considered some key parameters, such as pH conditions and solvent type, to assess their impact on redox potential prediction. Our GPR model predicted redox potentials with high accuracy across all datasets using minimal training data. The study provides powerful tools for molecule screening and design and delivers valuable guidance on designing training datasets for costly experiments.

25 ENERGY STORAGE↗

Chemistry Informed Machine Learning-Based Heat Capacity Prediction of Solid Mixed Oxides

Knowing heat capacity is crucial for modeling temperature changes with the absorption and release of heat and for calculating the thermal energy storage capacity of oxide mixtures with energy applications. The current prediction methods (ab initio simulations, computational thermodynamics, and the Neumann–Kopp rule) are computationally expensive, not fully generalizable, or inaccurate. Machine learning has the potential of being fast, accurate, and generalizable, but it has been scarcely used to predict mixture properties, particularly for mixed oxides. Here, we demonstrate a method for the generalizable prediction of heat capacity of solid oxide pseudobinary mixtures using heat capacity data obtained from computational thermodynamics and descriptors from ab initio databases. Further, models trained through this workflow achieved an error (mean absolute error of 0.43 J mol –1 K –1 ) lower than the uncertainty in differential scanning calorimetry measurements, and the workflow can be extended to predict other properties derived from the Gibbs free energy and for higher-order oxide mixtures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Results and lessons learned from accelerating radio frequency modeling using machine learning [slides]

The “advanced tokamak” reactor concept is a leading candidate for a steady state fusion pilot plant. An advanced tokamak (AT) sustains a majority of the required plasma current with effects resulting from maintenance of the peaked pressure at the device center. This current is augmented by auxiliary current drive sources. These auxiliary actuators may consist of neutral particle beams and/or radio frequency (RF) systems such as lower hybrid current drive (LHCD) and high harmonic fast wave (HHFW) current drive using radio and microwaves from antennas. The primary focus of this work is to develop models of RF current profile control suitable for use in integrated modeling frameworks and for real-time control in experiments. Direct physics models of RF current drive can be computationally intensive. In order to achieve predictive times appropriate for the thousands of calls needed in real-time control of experiments and for use in integrated models, we will apply modern machine learning (ML) techniques to accelerate these models and interpolate their results. To generate the fast and accurate models for use in control level algorithms and integrated modeling we need to replace present models with high dimensional interpolation of their results. We will perform additional simulations across a broader parameter range for EAST and other tokamaks in different physics regimes (Alcator C-Mod, DIII-D, WEST, CFETR, ARC, ITER) and combine them into a larger database for training and testing of the ML models. Further testing of the control level models with experimental current profile data from EAST and C-Mod tokamaks will provide additional confirmation of the control level model before integration in a tokamak control system or integrated modeling suite. ML will be used to optimize the selection of training data consisting of RF current driven at different values of density profile, temperature profile, plasma current, and wavenumber. ML will also be used to facilitate classification of current drive from these input data. The output of this effort will be a validated classifier capable of determining the current drive profiles for HHFW CD and LHCD on a mille-second timescale. This will provide a breakthrough capability enabling real-time control of RF driven current profiles in experiments including ITER ICRF and use integrated modeling frameworks requiring thousands of current profile calculations in discharge simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Shock Hugoniot calculations using on-the-fly machine learned force fields with ab initio accuracy

We present a framework for computing the shock Hugoniot using on-the-fly machine learned force field (MLFF) molecular dynamics simulations. In particular, we employ an MLFF model based on the kernel method and Bayesian linear regression to compute the free energy, atomic forces, and pressure, in conjunction with a linear regression model between the internal and free energies to compute the internal energy, with all training data generated from Kohn–Sham density functional theory (DFT). We verify the accuracy of the formalism by comparing the Hugoniot for carbon with recent Kohn–Sham DFT results in the literature. In so doing, we demonstrate that Kohn–Sham calculations for the Hugoniot can be accelerated by up to two orders of magnitude, while retaining ab initio accuracy. We apply this framework to calculate the Hugoniots of 14 materials in the FPEOS database, comprising 9 single elements and 5 compounds, between temperatures of 10 kK and 2 MK. We find good agreement with first principles results in the literature while providing tighter error bars. In addition, we confirm that the inter-element interaction in compounds decreases with temperature.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

ASCoT 3: Nonlinear Principal Components Analysis and Uncertainty Quantification in Early Concept Spacecraft Flight Software Cost Estimation

For mission planners and evaluators alike, value in cost models comes from a mean or median prediction, an understanding of the uncertainty on that prediction, and an understanding of model performance. Here we apply advanced statistical and machine learning methods to spacecraft flight software cost, effort, and SLOC estimation, and present the results in the latest version of the Analogy Software Cost Tool (ASCoT). We present in- and out-of-sample performance metrics for our models, each of which incorporate some amount of epistemic uncertainty. ASCoT, hosted on the One NASA Cost Engineering (ONCE) database via the Online NASA Space Estimation Tool (ONSET), was first showcased in 2016 as a number of analogy-based models and methods (kNN and Clustering) to support early project formulation. This ASCoT update improves upon the previous analogic methods by incorporating uncertainty in the data transformations. In particular, we use a Nonlinear Principal Components Analysis (NLPCA) to deal with ordinal data.

Robotic Spacecraft↗