Search NASA⌕ Search

SEARCH · Search NASA

Results for “Generative Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

A novel approach for large-scale wind energy potential assessment

Increasing wind energy generation is central to grid decarbonization, yet methods to estimate wind energy potential are not standardized, leading to inconsistencies and even skewed results. This study aims to improve the fidelity of wind energy potential estimates through an approach that integrates geospatial analysis and machine learning (i.e., Gaussian process regression). We demonstrate this approach to assess the spatial distribution of wind energy capacity potential in the Contiguous United States (CONUS). We find that the capacity-based power density ranges from 1.70 MW/km2 (25th percentile) to 3.88 MW/km2 (75th percentile) for existing wind farms in the CONUS. The value is lower in agricultural areas (2.73 ± 0.02 MW/km2, mean ± 95 % confidence interval) and higher in other land cover types (3.30 ± 0.03 MW/km2). Notably, advancements in turbine manufacturing could reduce power density in areas with lower wind speeds by adopting low specific-power turbines, but improve power density in areas with higher wind speeds (>8.35 m/s at 120m above the ground), highlighting opportunities for repowering existing wind farms. Wind energy potential is shaped by wind resource quality and is regionally characterized by land cover and physical conditions, revealing significant capacity potential in the Great Plains and Upper Texas. The results indicate that areas previously identified as hot spots using existing approaches (e.g., the west of the Rocky Mountains) may have a limited capacity potential due to low wind resource quality. Improvements in methodology and capacity potential estimates in this study could serve as a new basis for future energy systems analysis and planning.

Dai, Tao↗

Design of intrinsically disordered protein variants with diverse structural properties

Intrinsically disordered proteins (IDPs) perform a broad range of functions in biology, suggesting that the ability to design IDPs could help expand the repertoire of proteins with novel functions. Computational design of IDPs with specific conformational properties has, however, been difficult because of their substantial dynamics and structural complexity. We describe a general algorithm for designing IDPs with specific structural properties. We demonstrate the power of the algorithm by generating variants of naturally occurring IDPs that differ in compaction, long-range contacts, and propensity to phase separate. We experimentally tested and validated our designs and analyzed the sequence features that determine conformations. We show how our results are captured by a machine learning model, enabling us to speed up the algorithm. Our work expands the toolbox for computational protein design and will facilitate the design of proteins whose functions exploit the many properties afforded by protein disorder.

Science & Technology - Other Topics↗

Fast Machine Learning for Quantum Control of Microwave Qudits on Edge Hardware

Quantum optimal control is a promising approach to improve the accuracy of quantum gates, but it relies on complex algorithms to determine the best control settings. CPU or GPU-based approaches often have delays that are too long to be applied in practice. It is paramount to have systems with extremely low delays to quickly and with high fidelity adjust quantum hardware settings, where fidelity is defined as overlap with a target quantum state. Here, we utilize machine learning (ML) models to determine control-pulse parameters for preparing Selective Number-dependent Arbitrary Phase (SNAP) gates in microwave cavity qudits, which are multi-level quantum systems that serve as elementary computation units for quantum computing. The methodology involves data generation using classical optimization techniques, ML model development, design space exploration, and quantization for hardware implementation. Our results demonstrate the efficacy of the proposed approach, with optimized models achieving low gate trace infidelity near $10^{-3}$ and efficient utilization of programmable logic resources.

Sanders, Flor [Columbia U.]↗

Where IMERG Goes Next: Version 08 and Beyond

With the Version 07 (V07) Integrated Multi-satellitE Retrievals for GPM (IMERG) algorithm finalized and production initiated, the focus turns to enhancements for Version 08. These include innovations not included in V07 due to time constraints, plus issues revealed by the initial V07 products. One high priority is to evaluate and revise the schemes in V07 that rectify temporal artifacts caused by the time interpolation that fills the gaps between the various passive microwave (PMW) sensor overpasses. A second priority is to improve the homogeneity between the TRMM and GPM eras by characterizing differences between the two eras, determining the causes of these differences, and applying corrections as feasible, perhaps by enforcing spatial scale consistency (an overarching issue). Certainly, we must account for GPROF and the Combined Radar-Radiometer Algorithm converting to Machine Learning schemes in V08. Other priority topics include additional automated quality control for artifacts in the IR brightness temperatures and PMW precipitation fields, revisions to the specification algorithm for the probability of liquid precipitation, and accommodating new PMW sensors, which include the next generation of small-sats. We also consider the post-V08 landscape; the final GPM reprocessing will be restricted to fixing known code or algorithmic errors. Nonetheless, there are several data sources on the horizon to consider, including more small-sat PMW radiometers, AVHRR-based precipitation estimates (most useful in high latitudes), and the ISCCP-Next Generation and GEO-Ring projects that could provide easy access to multiple geosynchronous satellite channels and enable significantly improved algorithms compared to GEO-IR alone.

George J. Huffman↗

Verification, Validation, and Calibration Through a Causal Lens

While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

Aurora Detection From Nighttime Lights for Earth and Space Science Applications

This research leverages data from the Day/Night Band (DNB) of the Visible Infrared Imaging Radiometer (VIIRS) instrument onboard the Suomi National Polar-orbiting Partnership (S-NPP) satellite. We demonstrate the value of mining the VIIRS DNB for aurora and describe our use of unsupervised machine learning to create a binary mask for aurora occurrence. This mask can be used to flag aurora-contaminated observations for NASA's nighttime lights products for Earth science applications. The identification of auroral regions can also be used for Space Weather applications, for example, for comparison with aurora forecast model and with other satellite- or ground-based aurora observations. The DNB is a broadband channel that is sensitive to wavelengths from 500 to 900 nm, which covers most of the visible light spectrum, and as the name implies, captures light even at night with a sensitivity at the nanowatt level. This band is suitable for aurora observations since the light emitted by the aurora tends to be dominated by emissions from atomic oxygen, resulting in a greenish glow at a wavelength of 557.7 nm, especially at an altitude of 110 km. This study compares the global nighttime derived aurora regions for 17 and 18 March with the NOAA Space Weather Prediction Center's (SWPC) probability product for the St. Patrick's Day geomagnetic storm in 2015. VIIRS sensors are slated to be added to the next generation of polar-orbiting operational satellites. Our novel automated approach to aurora identification opens up an efficient way to leverage this unique data source.

Aurora↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

Applying deep learning methods to develop new models of molecular charge transfer, nonadiabatic dynamics, and nonlinear spectroscopy in the condensed phase

Photon- and field-induced charge transfer has central importance in the generation and storage of electricity, the novel properties of materials, photo-induced catalysis, and electro-optic activity (e.g., photovoltaic cells, fuel cells, and organic chromophores for use in optical fibers and light-emission diodes). These non-equilibrium electronic and chemical transformations are probed by ultrafast, nonlinear spectroscopies. Accurate simulations play a crucial role in our ability to understand, optimize, and control these transformations. This project applies modern deep learning and machine learning (ML) methods to dramatically improve models of electronic dynamics, electronic-nuclear dynamics, and spectroscopic measurements for improved simulations of chemistry in complex environments, far from equilibrium phenomena, and processes in extreme environments, such as materials exposed to strong or resonant fields. This project develops accurate neural net models that go beyond predictive capability to also provide new insight into the fundamental physics underlying electron and nuclear dynamics. To achieve its objectives, this project explores and develops customized versions of high-capacity deep learning algorithms/models. These techniques are developed with an emphasis on fundamental chemical insight, not just predictive accuracy, to assist the development of the next generation of quantum simulation methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structured Covariance Gaussian Networks for Orion Crew Module Aerodynamic Uncertainty Quantification

In this paper we propose a new approach for nonlinear regression and uncertainty quantification. The method is based on a pair of neural networks which parameterize mean and dense covariance functions of a multivariate Gaussian process, trained together to maximize the log-likelihood of observing the given data. The covariance matrix is made positive definite at every input by construction. We also propose a sampling approach that produces viable surrogate function realizations from the Gaussian process. We call the proposed model a Structured Covariance Gaussian Network (SCGN). We illustrate the use of SCGNs for learning an aerodynamic response surface with built-in uncertainty for the Orion crew module. We find that SCGN provides an efficient and systematic way to learn nonlinear functional relationships and dense covariances. We compare results to a baseline Gaussian process regressor and observe that the SCGN provides comparable uncertainty descriptions with improved scalability to dataset size. The sample functions generated by SCGN are fast to evaluate online and are therefore convenient for use in trajectory simulations. These results suggest that SCGN may be a viable computational method for aerodynamic uncertainty quantification.

machine learning↗

Structured Covariance Gaussian Networks for Orion Crew Module Aerodynamic Uncertainty Quantification

In this paper we propose a new approach for nonlinear regression and uncertainty quantification. The method is based on a pair of neural networks which parameterize mean and dense covariance functions of a multivariate Gaussian process, trained together to maximize the log-likelihood of observing the given data. The covariance matrix is made positive definite at every input by construction. We also propose a sampling approach that produces viable surrogate function realizations from the Gaussian process. We call the proposed model a Structured Covariance Gaussian Network (SCGN). We illustrate the use of SCGNs for learning an aerodynamic response surface with built-in uncertainty for the Orion crew module. We find that SCGN provides an efficient and systematic way to learn nonlinear functional relationships and dense covariances. We compare results to a baseline Gaussian process regressor and observe that the SCGN provides comparable uncertainty descriptions with improved scalability to dataset size. The sample functions generated by SCGN are fast to evaluate online and are therefore convenient for use in trajectory simulations. These results suggest that SCGN may be a viable computational method for aerodynamic uncertainty quantification.

machine learning↗

Bragg Spot Finder (BSF): a new machine-learning-aided approach to deal with spot finding for rapidly filtering diffraction pattern images

Macromolecular crystallography contributes significantly to understanding diseases and, more importantly, how to treat them by providing atomic resolution 3D structures of proteins. This is achieved by collecting X-ray diffraction images of protein crystals from important biological pathways. Spotfinders are used to detect the presence of crystals with usable data, and the spots from such crystals are the primary data used to solve the relevant structures. Having fast and accurate spot finding is essential, but recent advances in synchrotron beamlines used to generate X-ray diffraction images have brought us to the limits of what the best existing spotfinders can do. This bottleneck must be removed so spotfinder software can keep pace with the X-ray beamline hardware improvements and be able to see the weak or diffuse spots required to solve the most challenging problems encountered when working with diffraction images. In this paper, we first present Bragg Spot Detection (BSD), a large benchmark Bragg spot image dataset that contains 304 images with more than 66 000 spots. We then discuss the open source extensible U-Net-based spotfinder Bragg Spot Finder (BSF), with image pre-processing, a U-Net segmentation backbone, and post-processing that includes artifact removal and watershed segmentation. Finally, we perform experiments on the BSD benchmark and obtain results that are (in terms of accuracy) comparable to or better than those obtained with two popular spotfinder software packages ( Dozor and DIALS ), demonstrating that this is an appropriate framework to support future extensions and improvements.

36 MATERIALS SCIENCE↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally corrected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a surrogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [Illinois U., Chicago]↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Symbolic Execution Enhanced System Testing

We describe a testing technique that uses information computed by symbolic execution of a program unit to guide the generation of inputs to the system containing the unit, in such a way that the unit's, and hence the system's, coverage is increased. The symbolic execution computes unit constraints at run-time, along program paths obtained by system simulations. We use machine learning techniques treatment learning and function fitting to approximate the system input constraints that will lead to the satisfaction of the unit constraints. Execution of system input predictions either uncovers new code regions in the unit under analysis or provides information that can be used to improve the approximation. We have implemented the technique and we have demonstrated its effectiveness on several examples, including one from the aerospace domain.

Davies, Misty D.↗

Lunar Development Lab (LDL) Concept Leading to the First Human Lunar Outpost

The Lunar Development Lab (LDL) is a new concept to bring together academia, industry, non-profit organizations and NASA in an accelerator environment to generate new design solutions, technologies and architectures that will lead to the first human lunar outpost. By leveraging key partnerships in lunar science, mining, construction, chemical engineering and other key fields as well as making available rapid design, economic analysis, artificial intelligence (AI) and machine learning (ML) tools, significant progress can be made in a short amount of time. Therefore, the goal of LDL is to accelerate development and focus on economic solutions that can lead to sustainable and economical human lunar outpost.

Zuniga, Allison↗

Florida Water Resources: Assessing Coastal Resiliency Across Florida's Aquatic Preserves Response To Hurricane Forces

Intensifying weather events, sea level rise, and extensive coastal development in Southwestern Florida are escalating the need for Florida’s mangrove conservation. These mangroves are imperative for coastline stabilization, habitat provision for native species, and water quality management. Our partner, the Florida Department of Environmental Protection (FDEP), Office of Resilience and Coastal Protection is tasked with monitoring and conserving the Charlotte Harbor, Estero Bay, Rookery Bay, and Pinellas County Aquatic Preserves. We developed the Growth, Resilience, and Optical Vegetation Evaluator (GROVE) Google Earth Engine toolset for partners to determine mangrove forest extent through time, analyze mangrove forest health, and collect several water quality parameters within the preserves from January 2002–August 2022. The toolset provides easily accessible data from Landsat 7 Enhanced Thematic Mapper Plus (ETM+), Landsat 8 Operational Land Imager (OLI), Landsat 9 Operational Land Imager 2 (OLI-2), and the Shuttle Radar Topography Mission (SRTM). Using training datasets of known mangrove forest locations, we also established a machine learning approach to create mangrove extent maps. Maps from all four preserves indicated migration of mangrove forests inland as the greatest areas of change were transitional zones. Additionally, normalized difference vegetation index (NDVI), normalized difference turbidity index (NDTI), and chlorophyll-a maps were generated for the partners. This project provides decision makers with a useful tool for understanding temporal changes in Florida’s aquatic preserves, identifying areas of ecological stress, and providing actionable data to make informed plans for mangrove preservation.

Samuel Perrello↗