Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

An information-matching approach to optimal experimental design and active learning

The efficacy of mathematical models heavily depends on the quality of the training data, yet collecting sufficient data is often expensive and challenging. Many modeling applications require inferring parameters only as a means to predict other quantities of interest (QoI). Because models often contain many unidentifiable (sloppy) parameters, QoIs often depend on a relatively small number of parameter combinations. Therefore, we introduce an information-matching criterion based on the Fisher information matrix to select the most informative training data from a candidate pool. This method ensures that the selected data contain sufficient information to learn only those parameters that are needed to constrain downstream QoIs. It is formulated as a convex optimization problem, making it scalable to large models and datasets. Here, we demonstrate the effectiveness of this approach across various modeling problems in diverse scientific fields, including power systems and underwater acoustics. Finally, we use information-matching as a query function within an active learning (AL) loop for materials science applications. In all these applications, we find that a relatively small set of optimal training data can provide the necessary information for achieving precise predictions. These results are encouraging for diverse future applications, particularly AL in large machine-learning models.

Materials science↗

ES2Vec: Earth Science Metadata Suggestions and Analogical Reasoning

As the volume of text-based Earth science research grows, it is increasingly possible to discover latent relationships in the literature. However, traditional methodologies are restricted by limited computational capabilities and intractable problem spaces. Advancements in natural language processing (NLP) have allowed us to use an extensive Earth science corpus to create a domain-specific word vector model, Es2Vec, which we have used to surface latent relationships between Earth science concepts and generate improved keyword tags. Earth science metadata keyword assignment is a challenging problem. Dataset curators select appropriate keywords from the Global Change Master Directory (GCMD) set of keywords. The keywords an are integral part of the search and discovery of these datasets. Hence, the selection of keywords is crucial to increasing the discoverability of datasets. Utilizing machine learning techniques, we provide users with automated keyword suggestions to complement manual selection. We trained a machine learning model that leverages the semantic embedding ability of Word2Vec models to process abstracts and suggest relevant keywords. A user interface tool we built to assist data curators in the assignment of such keywords is also described.

word vectors↗

Variance Decomposition of MEDLI2 Reconstructed Heating Using Neural Networks

The Mars Entry, Descent, and Landing Instrumentation (MEDLI2) sensor suite collected data during entry of the Mars 2020 Perseverance rover into Mars’ atmosphere. This suite included a network of MEDLI2 Instrumented Sensor Plugs (MISPs). Each MISP was comprised of a cylinder made of Thermal Protection System (TPS) material with 1-3 embedded thermocouples (TCs), and it was flush mounted into the heatshield or backshell. Data from these in-depth TCs were used to reconstruct the aeroheating environment of the vehicle throughout entry. Surface heating was posed as an inverse problem, with the goal of estimating the surface heating by minimizing an objective function of the difference between MISP temperature measurements during flight and the temperature predictions derived from the Fully Implicit Ablation and Thermal response (FIAT) program. Given an aerothermal environment, FIAT calculates the material response and provides in-depth temperatures throughout the TPS material. To achieve the reverse, an internal tool called FIAT_Opt runs through multiple different environments until the output temperature at the TC depth closely matches the flight data. 95% confidence intervals on the reconstructed surface heating were obtained using Monte Carlo analysis, in which uncertainties in the thermocouple depth and the TPS material properties (e.g., density, thermal conductivity, heat capacity, emissivity) based on flight-lot material testing were included. A variance decomposition method using Sobol indices was employed to assess the sensitivity of the reconstructed peak heating to the TC placement and material property uncertainties. Variance decomposition was found to require tens of thousands of FIAT_Opt runs in order for the Sobol indices to converge. With a single FIAT_Opt run taking on the order of 40 minutes, the required number of computations would take months to complete, even if using multiple CPUs. To mitigate this problem, three machine learning models (ridge regression with cross-validation, random forest regression, and a deep neural network) were trained and tested using the 2000 Monte Carlo runs that were already completed. A subset of 1600 runs were used to train the model (i.e., training set), while the remaining 400 runs were used as the test set. The predictions from the deep neural network (DNN) on the test set showed nearly perfect agreement to the actual values computed with FIAT_Opt (R2 > 0.99). Using the DNN as a surrogate model, the variance decomposition using 50,000 runs was completed within minutes. The resulting Sobol indices showed that the reconstructed peak surface heating was most sensitive to the uncertainties in the thermal conductivity (ST = 0.37) and heat capacity (ST = 0.26). This method can be leveraged to provide requirements for material property measurements needed to improve the accuracy of surface heating prediction and ultimately lead to the reduction of design margins in the future. This presentation will include background on the MEDLI2 suite; the method used for inverse heating estimation; the way that material property uncertainties were accounted for using Monte Carlo analysis; a brief background on variance decomposition; the motivation for using machine learning in this context; how a neural network was trained on the data to enable variance decomposition in a fraction of the time; and the variance decomposition results for one of the MISPs.

Hannah Alpert↗

Presound: UAV Diagnostic System Enabled by Vibration-Based Machine Learning

A low-weight, inexpensive small unmanned aerial system (sUAS) that takes off, performs a mission, lands, and safely stows and recharges itself has myriad future applications ranging from agricultural imaging to last-mile package delivery. Likewise, Urban Air Mobility (UAM) systems will enable people to take air taxis from point to point in cities, rapidly moving commuters long distances without concern for road traffic and congestion. Fully electric aviation systems will be cleaner and quieter than ground transport. Cities could eliminate cars and buses, and convert roads to higher capacity bike and pedestrian throughways. Yet, for sUAS as well as UAM, system reliability and assurance is a limiting factor to deploying affordable autonomous flight systems. For this bright future of aviation to be realized, aircraft must be able to autonomously and accurately self-diagnose health issues both before takeoff and during flight. The GreenSight PreSound system is designed to identify defects on aircraft through intelligent analysis of vibration. It accomplishes this by measuring structural vibrations induced by the vehicle’s own propellers, and analyzing that data using a machine learning model that determines whether a defect is present. The PreSound system is designed to require no human oversight, and to operate across a wide array of vehicles through re-training of the model for each target aircraft. PreSound has been developed and seen limited early success using data collected from the GreenSight Dreamer sUAS, a 5lb quadrotor vehicle designed for aerial imaging applications. The final detection model, trained on data with props spinning at 50% throttle, achieves excellent performance with over 99% average accuracy in detecting blade damage using a single FFT vector input. It demonstrates the ability to generalize to new types of blade damage, correctly classifying a different type of blade damage with 98% accuracy. Full test pulses were classified with 100% accuracy, and in live testing, all sets of data during blade movement were classified accurately with over 95% confidence. When trained on in-flight data, the same model achieves an average accuracy of 85% in distinguishing between undamaged and blade-damaged states in flight. The authors believe that these accuracies show significant potential of this approach to expand unmanned flight safety, with significant potential benefits in accelerating Advanced Aerial Mobility (AAM) and UAM aviation applications.

UAS↗

Bridging Cloud and Edge Computing at NREL Using CONNECT: Cloud Optimized Networking for Next-Gen Edge Computing Technologies [Slides]

CONNECT is an innovative on-premise hardware and software solution that integrates edge and cloud computing infrastructure at NREL. Built on the AWS Greengrass middleware and leveraging the MQTT protocol, CONNECT enables real-time data streaming from IoT devices and gateways to both cloud and local services, empowering researchers to rapidly capture, analyze, and act upon edge-generated data while leveraging cloud capabilities. The platform addresses research infrastructure challenges by providing a pre-approved platform which is already configured with the correct networking and cybersecurity baselines thus eliminating procurement delays and enabling on-demand availability. CONNECT's hybrid architecture efficiently manages burstable workloads, allowing research teams to dynamically scale computational capacity, handle peak data loads, and reduce operational bottlenecks. Advanced capabilities include built-in GPU support for executing machine learning models which enables low-latency inference at the edge from models trained in the cloud. This architecture supports real-time analytics and filtering, providing a mechanism to allow only transmitting and processing high-value data. Cloud-based configuration management permits engineers to manage on-premise systems remotely, optimizing operational efficiency. By bridging edge and cloud computing, CONNECT provides NREL researchers with a flexible, scalable platform that accelerates scientific discovery while maintaining robust security and performance standards.

97 MATHEMATICS AND COMPUTING↗

High-resolution leaf area index maps generated from unoccupied aerial system, Teller Mile 27, Seward Peninsula, Alaska

Leaf area index (LAI), a measure of the amount of one-side leaf area per ground unit, is an important indicator of plant carbon, energy, and water cycle. In the heterogeneous Arctic landscapes, it has been challenging to accurately measure LAI across species and space needed for Earth system model validation. Here, we use multispectral unoccupied aerial systems (UASs) to scale up and map leaf area index (LAI) , in a low-Arctic tundra landscape on the Seward Peninsula, Alaska. We linked previous published LAI measurements with high-resolution, UAS-collected multispectral data collected over the region of Next Generation Ecosystem Experiments in the Arctic (NGEE Arctic)’s Teller Mile Maker 27 site in 2022 to develop random forest (RF) machine learning models to predict and map LAI. 100 RF models were developed to account for uncertainties in ground LAI plot measurements and process scaling. This dataset includes a raster (*.tif) map of the mean LAI value of the 100 RF models, a raster (*.tif) map of the standard deviation of the RF-modeled LAI data, and a user guide (*.pdf).

54 ENVIRONMENTAL SCIENCES↗

The macroevolution of filamentation morphology across the Saccharomycotina yeast subphylum

Saccharomycotina is a subphylum of ascomycete fungi with diverse asexual growth morphologies. Filamentous growth can comprise linear and branched budding cells that do not undergo cell separation, termed pseudohyphae, or tubular filaments with septa that perforate allowing movement of organelles, termed true hyphae. We integrated phenotypic, genomic, metabolic, and environmental data on isolation sources from 1051 species to examine the variation and evolutionary history of filamentation across Saccharomycotina and determine whether these data could predict filamentation types. We found that 63.37% of strains can form filaments; 6.56% true hyphae, 42.40% pseudohyphae, and 14.39% both true hyphae and pseudohyphae. The distributions of species that can produce true hyphae or filament were more strongly correlated with the yeast phylogeny than the distribution of species with pseudohyphae. Ancestral state reconstruction suggested that true hyphal and pseudohyphal morphologies evolved several times, that most yeast ancestors likely produced pseudohyphae or lacked filaments, and that the Saccharomycotina last common ancestor likely produced pseudohyphae but not true hyphae. Machine learning models trained on genomic and metabolic features predicted filament morphologies with ∼70% accuracy. Connecting the evolution of morphologies to their genomic, physiological, and ecological characteristics will enrich our understanding of how the diversity of lifestyles evolved in Saccharomycotina.

Saccharomycotina↗

A structured framework for predicting sustainable aviation fuel properties using liquid-phase FTIR and machine learning

Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.

Fourier transform infrared spectroscopy↗

Monitoring river flow status using low-cost wildlife camera and image segmentation artificial intelligence

Continuous measurement and monitoring of surface water coverage in non-perennial streams are essential for understanding the exchange fluxes between surface and subsurface waters under both inundated and non-inundated conditions. In this study, a wildlife camera photo-based framework was developed to monitor small stream water inundation, depth, discharge, and velocity. Two advanced machine learning models, YOLOv8 and Mask2Former, were utilized to efficiently analyze images captured by wildlife cameras. The accuracy of the framework was validated against on-site depth measurements at six sites in the Yakima River Basin, along with the gage height, discharge, and velocity data from four USGS sites. This approach facilitates long-term, continuous monitoring and quantification of river intermittency and water availability with high precision and low cost, thereby advancing river ecosystem research and management.

machine learning↗

The search for high-entropy fuel-cell catalysts using disorder descriptors

The transition to a hydrogen economy depends on efficient, affordable catalysts for fuel cells. Platinum—the industry standard for fuel-cell electrodes—is costly and scarce, highlighting the need for practical alternatives. High-entropy alloys offer vast compositional diversity and tunable properties that can mitigate these issues, yet their chemical complexity and configurational disorder have hindered rational discovery. Here, we introduce a data-driven framework that couples machine learning with first-principles disorder descriptors—including the entropy forming ability, disordered enthalpy-entropy descriptor, and electronic-structure similarity metrics to platinum—to predict alloy synthesizability and catalytic performance. These descriptors are applied for the first time in the context of fuel-cell catalyst discovery. The workflow rapidly screens more than 20 000 compositions and identifies several platinum-free candidates that are economically viable, readily scalable, and exhibit promising predicted activity. These results demonstrate that disorder descriptors are reliably predicted by machine learning models and can be effectively integrated into materials-discovery pipelines, accelerating innovation across complex compositional spaces.

fuel-cell catalysts↗

An Assessment of the Error Due to Computing Waste Isolation Pilot Plant Porosity Using the Porosity Response Surface Approach

The Waste Isolation Pilot Plant Performance Assessment (WIPP PA) must predict the likelihood that radionuclides will escape into the biosphere via mechanisms that depend on geohydraulic flow. Ideally, one would predict the geohydraulic flow using coupled geohydraulic and geomechanical simulations, but such coupled simulations are not computationally tractable. Instead, Sandia has historically used a look-up table of porosities for a given fluid pressure and time, called the porosity response surface, but this approach can introduce porosity errors because it largely ignores the porosity’s dependence on the past fluid pressure history. This report discusses efforts to quantify these porosity errors for both the legacy and new porosity response surfaces. Six hundred different fluid pressure histories were fed through the legacy/new geomechanical model and the legacy/new porosity response surface to generate six hundred porosity error histories. The error associated with the legacy porosity surface was substantial, while the error associated with the new porosity surface was typically small, except when fluid pressures exceeded the lithostatic pressure at the repository. In response to the errors at high pressures, a preliminary study of the WIPP PA’s sensitivity to these porosity errors was conducted. The study found that reducing the porosity errors at high pressures negligibly affected predictions of radionuclide releases. Finally, an initial machine-learned model for porosity was developed. This ML model significantly reduced the porosity error at high pressures, but sizable errors remained, so more development is necessary before coupling an ML model to the geohydraulic model.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Utilizing Earth Observations to Understand Landscape Patterns and Assist in Wildlife Management in Iona National Park, Angola

Following the end of the Angolan Civil War (1975-2002), human habitation in Iona National Park has grown exponentially, as has the livestock population. An ongoing drought beginning in 2017 has brought people, livestock, and wildlife into increasing competition for resources within the park. This study used Earth observation data, primarily Landsat and Sentinel imagery, to examine landscape trends to improve wildlife preservation approaches in Iona National Park, Angola. In collaboration with the NGO African Parks, we developed a robust land use and land cover (LULC) classification model using remote sensing data to augment sparse ground-based data in this arid land region. We used Google Earth Engine and a random forest classifier to map vegetation types, water bodies, and potential wildlife habitats. This analysis resulted in a high spatial resolution LULC time-series between 1984-2023, highlighting critical periods of socioecological change over the past 40 years. These results increased the partner’s ability to make scientifically grounded decisions about resource allocation and conservation priorities. This analysis supports the feasibility of applying remote sensing techniques coupled with machine learning models in dry regions, where standard survey methods are frequently limited by accessibility and resource availability. However, we identified limitations in ground-truth data and the difficulty of recognizing certain vegetation types in arid areas. Despite these limitations, the study demonstrated Earth observations' ability to transform wildlife management techniques in distant and data-scarce locations, providing a reproducible foundation for similar ecosystems around the world.

Emmanuel Aklie↗

Data-Driven Tailoring Optimization of Thermoset Polymers Using Ultrasonics and Machine Learning

Thermoset polymers are highly demanded for their structural robustness, thermal stability, and chemical resistance. Tailoring the properties of these polymers for high-performance applications is often preferred to designing brand-new polymers. However, the traditional destructive techniques used to characterize their properties as a function of manufacturing parameters are expensive and time-consuming. A novel non-destructive, data-driven method leveraging ultrasonics and machine learning techniques to tailor the properties of thermosets as a function of the manufacturing parameters is demonstrated. Thermoset epoxy samples with varying curing temperatures (15–40 °C) and curing agent amounts (±40%) were manufactured and tested. Their curing kinetics were monitored by determining the sound speed in the material in real time, while the longitudinal modulus of the samples was determined post-cure. Machine learning models were developed using a k-nearest neighbors algorithm. These models were implemented to predict the curing and final elastic properties using the manufacturing parameters, i.e., stoichiometry and curing temperature, and vice versa. Understanding and modeling how these parameters affect the cure kinetics and final properties will allow for efficient and reliable optimization of thermoset tailoring and manufacturing.

36 MATERIALS SCIENCE↗

Binding profiles for 961 Drosophila and C. elegans transcription factors reveal tissue-specific regulatory relationships

A catalog of transcription factor (TF) binding sites in the genome is critical for deciphering regulatory relationships. Here, we present the culmination of the efforts of the modENCODE (model organism Encyclopedia of DNA Elements) and modERN (model organism Encyclopedia of Regulatory Networks) consortia to systematically assay TF binding events in vivo in two major model organisms,Drosophila melanogaster(fly) andCaenorhabditis elegans(worm). These data sets comprise 605 TFs identifying 3.6 M sites in the fly and 356 TFs identifying 0.9 M sites in the worm, and represent the majority of the regulatory space in each genome. We demonstrate that TFs associate with chromatin in clusters termed “metapeaks,” that larger metapeaks have characteristics of high-occupancy target (HOT) regions, and that the importance of consensus sequence motifs bound by TFs depends on metapeak size and complexity. Combining ChIP-seq data with single-cell RNA-seq data in a machine-learning model identifies TFs with a prominent role in promoting target gene expression in specific cell types, even differentiating between parent–daughter cells during embryogenesis. These data are a rich resource for the community that should fuel and guide future investigations into TF function. To facilitate data accessibility and utility, all strains expressing green fluorescent protein (GFP)-tagged TFs are available at the stock centers for each organism. The chromatin immunoprecipitation sequencing data are available through the ENCODE Data Coordinating Center, GEO, and through a direct interface that provides rapid access to processed data sets and summary analyses, as well as widgets to probe the cell-type-specific TF–target relationships.

Biochemistry & Molecular Biology↗

Citizen science coupled with machine learning to quantify green-blue infrastructure cooling potential in Maricopa County, Arizona

Here, this study investigates the spatiotemporal cooling performance of green and blue infrastructure (GBI) in the Dobson Ranch urban neighborhood in Phoenix, Arizona. We leveraged citizen science near-surface (2 m) air temperature (Tair) measurements to train a highly accurate Tair predicting LightGBM machine learning model (R 2 : 0.986, MAE: 0.251 °C, RMSE: 0.585 °C). On June 16, 2024, the park area exhibited approximately 1 °C cooling effect (relative to the neighborhood mean) during both day and night. In contrast, the nearby artificial lake exhibited a stronger cooling effect of 2.4 °C during the day but a slight warming of 0.3 °C at night. At 00:00, locations 50 m downwind of the park were 0.3 °C warmer than the park, while locations 50 m upwind were 0.8 °C warmer. At 11:00, we observed that the downwind area is 0.8 °C cooler and the upwind area is 0.6 °C warmer—at the same 50 m distances relative to the park. We also observed 1 °C cooler and warmer effects respectively at the same 50 m downwind and upwind locations at 19:00 on June 17, 2024. Our data-driven analysis highlights potential limitations of car-traverse measurements, showing that failure to account for temporal variations during the traverse can lead to overestimation of Tair at night and underestimation during the day. Our analysis also showed only a weak correlation (coefficient: 0.48) between Landsat-derived land surface temperature (LST) and model predicted Tair at the time of the local Landsat overpass (∼11.00). This highlights the potential error of relying solely on LST for human thermal exposure analysis—particularly within the heterogenous built-environment.

54 ENVIRONMENTAL SCIENCES↗

Ionic Liquids for Direct Air Capture of CO 2 using Electric‐Field‐Mediated Moisture Gradient Process (Final Technical Report)

The final report provides executive summary, a list of publications, and information on training graduate students and postdoctoral researchers. We carried out computational and experimental research on understanding molecular-level mechanism of how CO 2 is absorbed in a solution containing ethylene glycol as the solvent and KOH as the salt in the presence of ionic liquids and under the influence of electric field. In doing so, we developed an automated high-throughput method which allowed us to measure the solubility of CO 2 in a large number of ionic liquids, considerably speeding up the CO 2 solubility measurement. We also demonstrated how varying the concentration of ionic liquids in ethylene glycol can result in a maximum in ionic conductivity. Reaction of CO 2 and subsequent release results in a 50% reduction when the process is operated at an ionic liquid-ethylene glycol concentration yielding maximum ionic conductivity amongst all the ionic liquid-ethylene glycol combinations studied as a part of this research. We utilized machine learning models to identify unique ionic liquid-solvent combinations with ionic conductivity much higher than that measured for ionic liquid-ethylene glycol combinations. We demonstrated that the rate of CO 2 reaction with KOH in ethylene glycol can be optimized with the type of ionic liquid and its concentration. Overall, the research led to publication of 10 peer-reviewed research articles and several presentations at national conferences. We are also in the process of developing additional manuscripts based on the research carried out as a part of this project. Two graduate students and two postdoctoral researchers were supported on the funding.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PyCMG-based Simulation of Volumetric Concrete Microstructure

Concrete is a complex, heterogeneous material with a microstructure composed of aggregates, cement paste, and pores spanning multiple length scales. Understanding this microstructure is critical for advancing the performance, durability, and modeling of concrete-based systems. While experimental imaging such as X-ray computed tomography (XCT) provides valuable insights, generating large datasets with detailed ground truth annotations is both costly and labor-intensive due to challenges in segmenting similar phases, such as aggregates and cement paste, that often share similar attenuation properties. To address this, we developed a pipeline to simulate realistic 3D concrete microstructures using the open-source Python package PyCMG. This simulation effort focuses on generating high-fidelity, annotated microstructures that can serve as training or benchmarking datasets for image analysis, segmentation algorithms, and machine learning models, particularly in scenarios where experimental data is scarce.

Ziabari, Amir [Oak Ridge National Laboratory; ORNL↗

Design of Materials with Alchemite

Machine learning models that establish the relationships between materials processing and properties can enable inverse design of materials through active learning. Alchemite is a commercial software that can perform inverse materials design on sparse data. Here we evaluate Alchemite’s performance on a dataset of shape memory alloys and a dataset of heat exchangers compared to baseline random forest models. Alchemite had higher accuracy when making predictions on sparse data and was more accurate or nearly as accurate as random forests on complete datasets while also quantifying uncertainty. The software was also used to suggest processing steps and design parameters to optimize properties and performance; however, physical validation of the suggested design parameters was beyond the scope of this work. Several useful design insights were gained about the impact of the design parameters on properties and performance including the importance of dopant choice and amount for shape memory alloys and the importance of height and weight on the thermal resistance of heat exchangers.

Machine learning↗