Search NASA⌕ Search

SEARCH · Search NASA

Results for “Generative Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55

Artificial Intelligence Transforming Post-Translational Modification Research

Post-Translational Modifications (PTMs) are covalent changes to amino acids that occur after protein synthesis, including covalent modifications on side chains and peptide backbones. Many PTMs profoundly impact cellular and molecular functions and structures, and their significance extends to evolutionary studies as well. In light of these implications, we have explored how artificial intelligence (AI) can be utilized in researching PTMs. Initially, rationales for adopting AI and its advantages in understanding the functions of PTMs are discussed. Then, various deep learning architectures and programs, including recent applications of language models, for predicting PTM sites on proteins and the regulatory functions of these PTMs are compared. Finally, our high-throughput PTM-data-generation pipeline, which formats data suitably for AI training and predictions is described. We hope this review illuminates areas where future AI models on PTMs can be improved, thereby contributing to the field of PTM bioengineering.

59 BASIC BIOLOGICAL SCIENCES↗

Hierarchical transfer learning: an agile and equitable strategy for machine-learning interatomic models

Machine-learned interatomic models are growing in popularity due to their ability to afford near quantum-accurate predictions for complex phenomena with orders-of-magnitude greater computational efficiency. However, these models struggle when applied to systems of many element types due to the approximately exponential increase in number of parameters that must be determined. To mitigate this challenge, we present a new hierarchical transfer learning approach that allows the fitting problem to be decomposed into smaller independent and reusable parameter blocks that enable development of explicitly chemically extensible ML-IAM. Application of this strategy is demonstrated for C and N mixtures under conditions ranging from nominally ambient to ~10,000 K and 200 GPa for compositions from 0 to 100% N. Ultimately, this strategy makes model generation for chemically complex systems more tractable and efficient, facilitates comprehensive model validation, and makes ML-IAM development for problems of this nature more accessible to users with limited access to extreme computing infrastructure.

Lindsey, Rebecca K. [Univ. of Michigan, Ann Arbor,↗

TransPlatformer

We propose TransPlatformer for translating toxicogenomics from one platform to another. Transcriptomic profiling has evolved through multiple generations of technology, from microarrays (e.g., Affymetrix, CodeLink) to more recent high-throughput sequencing and targeted panels such as S1500+. Microarrays, which dominated gene expression studies in the early 2000s, provided affordable and high-throughput transcript quantification but suffered from cross-hybridization issues and limited dynamic range . RNA-Seq, introduced in the late 2000s, revolutionized transcriptomics by enabling unbiased and comprehensive gene expression analysis, albeit at higher costs and computational demands . Despite advances, many studies rely on historical microarray data, necessitating the translation of legacy data into modern platforms to ensure continuity and comparability. This translation is complicated by factors such as platform-specific probe design, differences in transcript coverage, and batch effects . Existing methods for cross-platform mapping include statistical normalization, machine learning models, and biological anchoring approaches. The ability to translate transcriptomic data between platforms has broad implications, including enhanced meta-analyses, improved toxicological modeling, and better integration of historical datasets with contemporary research. TransPlatformer seeks to contribute to this effort by evaluating translation methodologies and proposing novel strategies to improve cross-platform gene expression harmonization. In this repository there are code examples for TransPlatformer implementation

Cong, Guojing↗

Satellite Embedding-Based Population Imputation for Areas with Missing Building Footprint Data: A Computer Vision-Based Approach

High-resolution population modeling is important for supporting effective decision-making across diverse sectors. LandScan Mosaic generates population estimates at the level of individual buildings and aggregates them to 3 arc-second grids, and this approach performs well in regions where building footprint data are comprehensive and reliable. However, large portions of the globe still suffer from incomplete, sparse, or entirely missing building stock datasets, creating a structural limitation for strictly building-based population models. To address this research gap, this study proposes a computer vision-based framework that employs Google Earth Engine satellite embeddings and UNet, which allows us to directly impute grid-level population estimates in building-data-deficient areas. Applied to Taiwan as a case study, the framework achieved strong predictive performance with R$^{2}$ of 0.89, RMSE of 18.70, and MAE of 8.41, outperforming traditional machine learning approaches. Notably, the proposed framework effectively addressed building false-positive errors inherent in Global Human Settlement Layer (GHSL) data, correctly identifying uninhabited areas that were erroneously classified as populated. The framework also offers significant advantages for global population mapping, particularly in terms of scalability and temporal consistency, thereby extending the coverage and accuracy of high-resolution population products in data-scarce regions worldwide. Urban planners, decision makers, and related stakeholders can obtain granular population distributions to support more accurate and targeted infrastructure investment, service delivery, resource allocation, and risk assessment decisions.

97 MATHEMATICS AND COMPUTING↗

Model Generation to Support Model-Based Testing Applied on NASA DAT - An Experience Report

Model-based Testing (MBT), where a model of the system under tests (SUT) behavior is used to automatically generate executable test cases, is a promising and versatile testing technology. Nevertheless, adoption of MBT technologies in industry is slow and many testing tasks are performed via manually created executable test cases (i.e. test programs such as JUnit). In order to adopt MBT, testers must learn how to construct models and use these models to generate test cases, which might be a hurdle. An interesting observation in our previous work is that the existing manually created test cases often provided invaluable insights for the manual creation of the testing models of the system. In this paper we present an approach that allows the tester to first create and debug a set of test cases. When the tester is happy with the test cases, the next step is to automatically generate a model from the test cases. The generated model is derived from the test cases, which are actions that the system can perform (e.g. a button clicks) and their expected outputs in form of assert statements (e.g. assert data entered). The model is a Finite State Machine (FSM) model that can be employed with little or no manual changes to generate additional test cases for the SUT. We successfully applied the approach in a feasibility study to the NASA Data Access Toolkit (DAT), which is a web-based GUI. One compelling finding is that the test cases that were generated from the automatically generated models were able to detect issues that were not detected by the original set of manually created test cases. We present the findings from the case study and discuss best practices for incorporating model generation techniques into an existing testing process.

State Machines↗

Unveiling the Hidden Evolution of Crystal Defects and Disorder in Energy Materials

Control of point defects and disorder in functional thin films and 2D materials is critical to realizing their full potential in applications ranging from energy storage to advanced electronics. However, these phenomena are often poorly understood, difficult to characterize, and challenging to direct with precision. This presentation explores emerging multi-modal computer vision to decipher and predict order in materials across multiple length scales in the electron microscope, from the atomic to the nanoscale. By fusing data from diverse sources, these powerful models provide unprecedented insights into materials' lifecycles, enabling the control of defects and their associated properties at a fundamental level. This capability promises to transform materials design and accelerate the development of next-generation technologies.

97 MATHEMATICS AND COMPUTING↗

Weighted Composition Operators for Learning Nonlinear Dynamics

Operator theoretic methods in dynamical system have been dominated by the use of Koopman operators and their continuous time counterparts, such as Koopman Generators and Liouville Operators. The advantage gained from their use primarily stems from the ability to extract subspaces and eigenfunctions within a space of observables that are invariant with respect to the Koopman operator over that space. When this occurs, a dynamic mode decomposition of the systems state provides a linear model for the dynamical system. Not all Koopman operators have eigenfunctions that may be exploited in this manner. However, the framework can still be leveraged for approximations using other operators. In this setting, we present a different operator for the study of dynamical systems, the weighted composition operator. These operators are compact for a wide range of dynamics and spaces, and through their interactions with occupation kernels and vector valued kernels, they admit an estimation of the underlying dynamics. Here, this manuscript presents a new algorithm for the data driven study of dynamical systems from data, and also provides two numerical experiments where convergence is achieved as a proof of concept.

97 MATHEMATICS AND COMPUTING↗

REIMAGINING HEAT EXCHANGERS FOR NEXT GENERATION ENVIRONMENTAL SYSTEMS

Air-to-refrigerant heat exchangers (HXs) are essential components in space conditioning, refrigeration, and power systems, and recent efforts have focused on making these devices more compact, reducing refrigerant charge and lowering manufacturing costs. Historically, HX innovation has been limited by available computational resources, design tools, and manufacturing constraints. The best available technologies utilize tube-fin and micro- or macro-channel tubes with fins, which are not necessarily the optimal designs achievable with current technology. In this paper, we highlight the latest advancements in air-to-refrigerant HXs, specifically emphasizing innovations achieved through shape and topology optimization. A multi-scale design optimization approach is introduced, alongside similar methods in literature, which enable highly sophisticated shape-optimized tube designs with more than 50% reduction in size and 25% reduction in refrigerant charge, essential for A3 refrigerant charge limit compliance. The frameworks integrate traditional heat and mass transfer science with state-of-the-art machine learning, genetic algorithms, and adjoint algorithms to create novel designs. While many of these innovative designs may not be manufacturable using conventional methods, they allow us explore the boundaries of what is possible. These novel air-to-refrigerant HXs are key enablers for ultra-low-refrigerant charge heat pump and refrigeration systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The high explosives & affected targets (HEAT) dataset

Artificial Intelligence (AI) surrogate models offer a computationally efficient alternative to full-physics simulations, yet no existing datasets are publicly available for training, testing, and validation of machine learning models of the dynamics of high-explosive driven shocks through multiple materials. Shock propagation through materials is a computationally challenging problem because simulations must include material-specific equations of state (EOS) along with descriptions of other physical processes such as plastic deformation, phase change, damage processes, fluid instabilities, and multi-material interactions. Shocks are typically initiated by high-velocity impacts or explosive loading. The latter case necessitates the addition of models of reactive materials to represent high-explosive (HE) detonation. Here, to address the lack of an expansive dataset for multi-material shock propagation in the AI/ML community, we present the High-Explosives and Affected Targets (HEAT) Dataset. HEAT is a physics-rich collection of two-dimensional, cylindrically symmetric, simulations generated using an Eulerian, multi-material, shock-propagation code developed at Los Alamos National Laboratory. The dataset includes two partitions: (1) the expanding shock-cylinder (CYL) simulations, Figs. 1, and (2) the Perturbed Layered Interface (PLI) simulations, Fig. 2. Entries in both partitions consist of time series of arrays of thermodynamic fields (pressure, density, and temperature), kinematic fields (position and velocity), and additional fields that depend on thermodynamic and/or kinematic fields (e.g., material stress). Materials in the CYL partition include solids (aluminium, copper, depleted uranium, stainless steel, tantalum, and a generic polymer), a liquid (water), gases (air, nitrogen), and a generic detonating material (high explosive, HE). The PLI partition spans a highly varying geometry but consists of fixed materials across entries: Copper, aluminium, stainless steel, generic polymer, and generic HE. HEAT captures critical phenomena such as momentum transfer, shock propagation, plastic deformation, and thermal effects, making HEAT a valuable benchmark for development of AI/ML emulation of multi-material shock propagation.

36 MATERIALS SCIENCE↗

Integrated Approach to Ancillary PV Component Reliability Assessment (Final Report)

In this project, we have established a nondestructive, generalized methodology that (1) fuses rich field data with advanced ML for proactive reliability forecasting, (2) dramatically reduces experimental iterations via synthetic dataset generation, and (3) achieves unprecedented regression precision in both anomaly detection and component-level degradation assessment—paving the way for truly predictive maintenance of grid-tied PV inverters under diverse outdoor conditions.

14 SOLAR ENERGY↗

Temporally-consistent koopman autoencoders for forecasting dynamical systems

Absence of sufficiently high-quality data often poses a key challenge in data-driven modeling of high-dimensional spatio-temporal dynamical systems. Koopman Autoencoders (KAEs) harness the expressivity of deep neural networks (DNNs), the dimension reduction capabilities of autoencoders, and the spectral properties of the Koopman operator to learn a reduced-order feature space with simpler, linear dynamics. However, the effectiveness of KAEs is hindered by limited and noisy training datasets, leading to poor generalizability. To address this, we introduce the Temporally-Consistent Koopman Autoencoder (tcKAE), designed to generate accurate long-term predictions even with limited and noisy training data. This is achieved through a consistency regularization term that enforces prediction coherence across different time steps, thus enhancing the robustness and generalizability of tcKAE over existing models. We provide analytical justification for this approach based on Koopman spectral theory and empirically demonstrate tcKAE’s superior performance over state-of-the-art KAE models across a variety of test cases, including simple pendulum oscillations, kinetic plasma, and fluid flow data.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Locomotion training of legged robots using hybrid machine learning techniques

In this study artificial neural networks and fuzzy logic are used to control the jumping behavior of a three-link uniped robot. The biped locomotion control problem is an increment of the uniped locomotion control. Study of legged locomotion dynamics indicates that a hierarchical controller is required to control the behavior of a legged robot. A structured control strategy is suggested which includes navigator, motion planner, biped coordinator and uniped controllers. A three-link uniped robot simulation is developed to be used as the plant. Neurocontrollers were trained both online and offline. In the case of on-line training, a reinforcement learning technique was used to train the neurocontroller to make the robot jump to a specified height. After several hundred iterations of training, the plant output achieved an accuracy of 7.4%. However, when jump distance and body angular momentum were also included in the control objectives, training time became impractically long. In the case of off-line training, a three-layered backpropagation (BP) network was first used with three inputs, three outputs and 15 to 40 hidden nodes. Pre-generated data were presented to the network with a learning rate as low as 0.003 in order to reach convergence. The low learning rate required for convergence resulted in a very slow training process which took weeks to learn 460 examples. After training, performance of the neurocontroller was rather poor. Consequently, the BP network was replaced by a Cerebeller Model Articulation Controller (CMAC) network. Subsequent experiments described in this document show that the CMAC network is more suitable to the solution of uniped locomotion control problems in terms of both learning efficiency and performance. A new approach is introduced in this report, viz., a self-organizing multiagent cerebeller model for fuzzy-neural control of uniped locomotion is suggested to improve training efficiency. This is currently being evaluated for a possible patent by NASA, Johnson Space Center. An alternative modular approach is also developed which uses separate controllers for each stage of the running stride. A self-organizing fuzzy-neural controller controls the height, distance and angular momentum of the stride. A CMAC-based controller controls the movement of the leg from the time the foot leaves the ground to the time of landing. Because the leg joints are controlled at each time step during flight, movement is smooth and obstacles can be avoided. Initial results indicate that this approach can yield fast, accurate results.

Simon, William E.↗

Agent-Based, Bottom-Up Medium- and Heavy-duty Electric Vehicle Economics, Operation, Charging and Adoption (Research Performance Final Report)

This is the research performance final report for the project entitled: Agent-Based, Bottom-Up Medium- and Heavy-duty Electric Vehicle Economics, Operation, Charging and Adoption This project was able to achieve the DOE’s goals of developing new modeling tools to understand MDHD vehicle operation and adoption. The first modeling tool is a fleet-level techno-economic analysis model capable of estimating energy use and associated environmental and cost impacts for electrified and conventional vehicles of any MDHD vocation, using real-world cost and operations data, including approaches to optimizing schedules for charging and/or vehicle dispatch. The second modeling tool is a system-level, bottom-up, agent-based adoption model capable of generating geographically-resolved estimates of market projections for MDHD vehicles and charging infrastructure. These tools will be developed and published to serve dual purposes as analysis tools for researchers, and decision-support tools for decision makers within the MDHD system.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A model to assess Zircaloy’s mechanical property changes following a transient beyond critical heat flux

Maintaining the integrity of nuclear fuel rods is essential for ensuring public health and safety in nuclear power generation. During reactor operation, this integrity is confirmed by demonstrating compliance with established regulatory acceptance criteria. For moderate-frequency events, such as limiting transients and anticipated operational occurrences (AOOs), the current fuel integrity criterion is based on preventing boiling transition. This criterion assumes that prevention of boiling transition will prevent excessive cladding heating and, thus, fuel failure during normal operations. While conservative, this approach places significant constraints on core design, fuel cycle economics, and a plant’s ability to perform major power uprates, leading to suboptimal fuel utilization and inefficient carbon-free energy production. A more efficient approach could be achieved by revising the failure criterion to a material-specific limit rather than strictly preventing the boiling transition, since boiling transition per se is not a cause of fuel cladding failure. Here, as a result, a new licensing framework based on material properties, termed time-at-temperature (t@T), is needed. This approach would allow for brief periods of post–critical heat flux operation during an AOO without compromising safety. Implementing the t@T licensing strategy requires a robust technical foundation in material properties, which must be established through comprehensive data collection on both unirradiated and irradiated fuel and cladding materials. This foundation would enable the development of a safety basis that ensures safe operation while providing greater flexibility and efficiency for reactor operation. This paper documents a thorough review of the available data to establish a baseline knowledge that can inform the development of cladding mechanical models, as well as identify experimental data gaps that need to be addressed in future research. Machine learning and data informatics were utilized to extract the importance of parameters on the t@T parameter. Industry tools were used to perform baseline analyses to define the relevant transient conditions for data analysis. The subsequent review successfully identified applicable experimental data, as well as sufficient data to evaluate changes in cladding mechanical properties following an AOO transient. Rather than developing new models, this work coupled existing irradiation annealing and recrystallization models to calculate changes in hardness, yield stress, and ultimate tensile stress following an AOO event. The findings from this review were summarized to highlight the experimental data needs required to fill remaining gaps and support the development of future t@T licensing methodologies.

Cladding performance↗

Extending High-Level Synthesis with AI/ML Methods

Artificial Intelligence (AI) and Machine Learning (ML) methods provide significant opportunities of improving quality of results when performing high-level synthesis (HLS). For example, they can be used to model and predict metrics of the final design (e.g., area, considering aspects such as interconnect overhead for different device technologies), facilitating exploration when searching for the best design trade-offs. They can also enable identifying hidden correlations across the various phases of the synthesis and the various optimizations performed, identifying the most effective pipelines. Finally, in more general terms, bio-inspired heuristic algorithms can improve the design space exploration for the synthesis process in terms of time and quality of the result. This paper discusses opportunities and challenges to augment HLS with AI/ML using as example flow the SODA Synthesizer, an open-source hardware generation toolchain which includes SODA-OPT, a hardware/software partitioning and pre-optimization tool developed with the MLIR framework, and PandA-Bambu, a state-of-the art HLS tool. SODA interfaces with OpenROAD to provide a complete end-to-end toolchain.

artificial intelligence↗

A New Generation of Intelligent Trainable Tools for Analyzing Large Scientific Image Databases

In a variety of scientific disciplines two-dimensional digital image data is now relied on as a basic component of routine scientific investigation. The proliferation of image acquisition hardware such as multi-spectral remote-sensing platforms, medical imaging sensors, and high-resolution cameras have led to the widespread use of image data in fields such as atmospheric studies, planetary geology, ecology, agriculture, glacielogy, forestry, astronomy, diagnostic medicine, to name but a few.

machine↗

Robust Semantic Mapping and Localization on a Free-Flying Robot in Microgravity

We propose a system that uses semantic object detections to localize a microgravity free-flyer. Many applications require absolute localization in a known reference frame, such as the execution of waypoint trajectories defined by human operators. Classical geometric methods build a map of point features, which may not be able to be associated after lighting or environmental changes. By contrast, semantics remain invariant to changes up to the robustness of the detection algorithm and motion of the semantic objects. In this work, we describe our approaches for both offline semantic map generation as well as online localization against a semantic map, intended to run in real-time on the robot. We additionally demonstrate how our semantic localizer outperforms image-feature matching in some cases, and show the robustness of the algorithm to environmental changes. Crucially, we show in our experiments that when semantics are used to supplement point features, localization is always improved. To our knowledge, these experiments demonstrate the first use of learned semantics for localization on a free-flying robot in microgravity.

Localization↗

Machine learning-enabled computer vision for plant phenotyping: a primer on AI/ML and a case study on stomatal patterning

Abstract Artificial intelligence and machine learning (AI/ML) can be used to automatically analyze large image datasets. One valuable application of this approach is estimation of plant trait data contained within images. Here we review 39 papers that describe the development and/or application of such models for estimation of stomatal traits from epidermal micrographs. In doing so, we hope to provide plant biologists with a foundational understanding of AI/ML and summarize the current capabilities and limitations of published tools. While most models show human-level performance for stomatal density (SD) quantification at superhuman speed, they are often likely to be limited in how broadly they can be applied across phenotypic diversity associated with genetic, environmental, or developmental variation. Other models can make predictions across greater phenotypic diversity and/or additional stomatal/epidermal traits, but require significantly greater time investment to generate ground-truth data. We discuss the challenges and opportunities presented by AI/ML-enabled computer vision analysis, and make recommendations for future work to advance accelerated stomatal phenotyping.

Plant Sciences↗