Search NASASearch

SEARCH · Search NASA

Results for “learning (artificial intelligence)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Quantum-Inspired Bayesian Sampling for Uncertainty Quantification and Machine Learning (Final Technical Report)

With increasing simulation and measurement data, machine learning and artificial intelligence have been widely used in computational decision-making of complex engineering systems. The resulting tools, such as uncertainty quantification solvers, reinforcement learning, and physics-informed machine learning, have achieved great success in critical DOE tasks such as material discovery and design, energy system modeling and control, and numerical weather and climate prediction. A core topic in scientific machine learning and artificial intelligence is Bayesian inference: given an observed data set, people want to estimate the posterior distribution of a (possibly large) number of hidden parameters. Due to the flexibility and weak assumptions, Bayesian sampling has been the mainstream Bayesian inference solvers despite the rapid progress of approximate Bayesian inference. Classical Bayesian sampling methods such as Markov-chain Monte Carlo suffer from a low-acceptance rate due to the random walk nature, therefore state-of-the-art techniques use Hamiltonian Monte Carlo and its variants to efficiently draw posterior samples in a high dimension. The key idea of Hamiltonian Monte Carlo and its variants is to simulate the Hamiltonian dynamics of a classical particle with a fixed mass, and their performance significantly degrades when the posterior distribution is highly spiky or has multiple modes. Leveraging the idea of quantum physics, this project has investigated new theory, algorithms and applications of Bayesian inference (especially Bayesian sampling). The main results include: (1) novel quantum-inspired Bayesian sampling methods that can lead to better accuracy for challenging multi-modal or spiky distributions, (2) more scalable machine learning framework leveraging tensor-compressed Bayesian inference, and (3) Bayesian and sampling approaches for verifying the robustness of continuous and binary neural networks.

97 MATHEMATICS AND COMPUTING

Leafweb: Leaf Gas Exchange and Pulse-Amplitude Modulated Fluorometry for C4 Species, June 2026 Release

This dataset contains leaf gas exchange and Pulse-Amplitude Modulated (PAM) fluorometry for 98 C4 species. The C4 photosynthetic pathway employs specialized CO2 concentration mechanisms and Kranz anatomy to enrich CO2 concentration around Rubisco, the enzyme that catalyzes carbon fixation in the Calvin-Benson cycle to suppress photorespiration and increase the use efficiencies of light, nitrogen, and water as compared to the C3 photosynthetic pathways. Large-scale C4 photosynthetic datasets are relatively scarce, which has affected C4 photosynthesis research. To improve C4 photosynthetic data availability, Leafweb organized an effort to systematically collect, compile, standardize, and organize measurements of leaf gas exchange and/or Pulse-Amplitude Modulated (PAM) fluorometry of C4 species. This derived a C4 photosynthetic dataset containing measurements made by independent researchers in multiple countries in various environments (field, garden, or greenhouse). It covers three biochemical subtypes – the nicotinamide adenine dinucleotide phosphate-malic enzyme (NADP-ME), nicotinamide adenine dinucleotide-malic enzyme (NAD-ME), and phosphoenolpyruvate carboxykinase (PEP-CK) subtypes. This dataset is useful for using Artificial Intelligence / Machine Learning and mechanistic models to study C4 photosynthesis and compare across different biochemical subtypes. This dataset contains 3 compressed (*.zip) folders containing 1,892 data files in comma-separate values (*.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma-separate values (*.csv) format and a user guide in PDF (*.pdf) format.

Zhou, Haoran [Tianjin University, China]

SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions

Instruction finetuning is a popular paradigm to align large language models (LLM) with human intent. Despite its popularity, this idea is less explored in improving the LLMs to align existing foundation models with scientific disciplines, concepts and goals. In this work, we present SciTune as a tuning framework to improve the ability of LLMs to follow scientific multimodal instructions. To test our methodology, we use a human-generated scientific instruction tuning dataset and train a large multimodal model LLaMA-SciTune that connects a vision encoder and LLM for science-focused visual and language understanding. LLaMA-SciTune significantly outperforms the state-of-the-art models in the generated figure types and captions in multiple scientific multimodal benchmarks. In comparison to the models that are fine-tuned with machine generated data only, LLaMA-SciTune surpasses human performance on average and in many sub-categories on the ScienceQA benchmark.

• Artificial intelligence (AI) / machine learning

Universal Fourier Attack for Time Series

A wide variety of adversarial attacks have been proposed and explored using image and audio data. These attacks are notoriously easy to generate digitally when the attacker can directly manipulate the input to a model, but are much more difficult to implement in the real world. In this paper we present a universal, time invariant attack for general time series data such that the attack has a frequency spectrum primarily composed of the frequencies present in the original data. The universality of the attack makes it fast and easy to implement as no computation is required to add it to an input, while time invariance is useful for real world deployment. Additionally, the frequency constraint ensures the attack can withstand filtering defenses. We demonstrate the effectiveness of the attack on two different classification tasks through both digital and real world experiments, and show that the attack is robust against common transform-and-compare defense pipelines.

97 MATHEMATICS AND COMPUTING

Leveraging generative AI for urban digital twins: a scoping review on the autonomous generation of urban data, scenarios, designs, and 3D city models for smart city advancement

The digital transformation of modern cities by integrating advanced information, communication, and computing technologies has marked the epoch of data-driven smart city applications for efficient and sustainable urban management. Despite their effectiveness, these applications often rely on massive amounts of high-dimensional and multi-domain data for monitoring and characterizing different urban sub-systems, presenting challenges in application areas that are limited by data quality and availability, as well as costly efforts for generating urban scenarios and design alternatives. As an emerging research area in deep learning, Generative Artificial Intelligence (GenAI) models have demonstrated their unique values in content generation. This paper aims to explore the innovative integration of GenAI techniques and urban digital twins to address challenges in the planning and management of built environments with focuses on various urban sub-systems, such as transportation, energy, water, and building and infrastructure. The survey starts with the introduction of cutting-edge generative AI models, such as the Generative Adversarial Networks (GAN), Variational Autoencoders (VAEs), Generative Pre-trained Transformer (GPT), followed by a scoping review of the existing urban science applications that leverage the intelligent and autonomous capability of these techniques to facilitate the research, operations, and management of critical urban subsystems, as well as the holistic planning and design of the built environment. Based on the review, we discuss potential opportunities and technical strategies that integrate GenAI models into the next-generation urban digital twins for more intelligent, scalable, and automated smart city development and management.

3D city modeling

AI/ML-assisted Design of Phosphate Glass and Ceramic Nuclear Waste Forms

Borosilicate glass is the widely accepted waste form for immobilization of high and medium level nuclear wastes. Advances in nuclear energies and new reactor designs require the development of new waste forms. For example, wastes from molten salt reactors and reprocessing of nuclear fuels lead to salt-based wastes that are difficult to be immobilized by conventional borosilicate glasses due to limited solubility and waste loading. In designing new waste forms, machine learning (ML) and artificial intelligence (AI) based approaches are much needed and can be beneficial in enabling a more efficient design in large parameter spaces as compared to traditional Edisonian trial-and-error approaches. Here, we report in this paper the rationale and latest progress of our ML/AI-based design of phosphate-based waste forms.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Accelerated Selectrion of Optimal Perovskite Alloys for Solar PV using a Combined Quantum and Machine Learning Hierachiral Approach

The project aims to: (i) accelerate the discovery of “Missing HP alloys” by combining quantum mechanics and artificial intelligence machine learning approaches, and (ii) analyze the stabilities of candidate alloys, including those that do not pass selection filters (and are thus expected to degrade over time) to decipher the nature of the instabilities to guide the development of durable solar cell materials. Successful candidates will be subjected to validation experiments at NREL's state-of-the-art facilities. The discoveries this effort will provide will be directly testable and implementable and will greatly impact U.S. progress in HP PV as they will provide clear direction and motivation for experimental studies including specific material synthetic targets, device optimization, and device stability protocols. A key advantage of this effort is the feedback and guidance provided by the Industry Collaborative Work Group that we established to coordinate academic and national lab research with industry needs. The proposed work will provide a basis for and direct the development of robust and reliable HP PV. It will also provide a timely, valuable and extensive roadmap to the experimental HP PV community to enable it to focus its efforts on improving and fine-tuning promising HP compositions that this effort predicts will likely be the best performers rather than wandering in the vast chemical space for decades spending enormous resources mostly evaluating unpromising candidate materials.

14 SOLAR ENERGY

Moving beyond post hoc explainable artificial intelligence: a perspective paper on lessons learned from dynamical climate modeling

AI models are criticized as being black boxes, potentially subjecting climate science to greater uncertainty. Explainable artificial intelligence (XAI) has been proposed to probe AI models and increase trust. In this review and perspective paper, we suggest that, in addition to using XAI methods, AI researchers in climate science can learn from past successes in the development of physics-based dynamical climate models. Dynamical models are complex but have gained trust because their successes and failures can sometimes be attributed to specific components or sub-models, such as when model bias is explained by pointing to a particular parameterization. We propose three types of understanding as a basis to evaluate trust in dynamical and AI models alike: (1) instrumental understanding, which is obtained when a model has passed a functional test; (2) statistical understanding, obtained when researchers can make sense of the modeling results using statistical techniques to identify input–output relationships; and (3) component-level understanding, which refers to modelers' ability to point to specific model components or parts in the model architecture as the culprit for erratic model behaviors or as the crucial reason why the model functions well. We demonstrate how component-level understanding has been sought and achieved via climate model intercomparison projects over the past several decades. Such component-level understanding routinely leads to model improvements and may also serve as a template for thinking about AI-driven climate science. Currently, XAI methods can help explain the behaviors of AI models by focusing on the mapping between input and output, thereby increasing the statistical understanding of AI models. Yet, to further increase our understanding of AI models, we will have to build AI models that have interpretable components amenable to component-level understanding. We give recent examples from the AI climate science literature to highlight some recent, albeit limited, successes in achieving component-level understanding and thereby explaining model behavior. The merit of such interpretable AI models is that they serve as a stronger basis for trust in climate modeling and, by extension, downstream uses of climate model data.

54 ENVIRONMENTAL SCIENCES

Intelligent experiments through real-time AI: Fast Data Processing and Autonomous Detector Control for sPHENIX and future EIC detectors

This R&D project, initiated by the DOE Nuclear Physics AI-Machine Learning initiative in 2022, leverages AI to address data processing challenges in high-energy nuclear experiments (RHIC, LHC, and future EIC). Our focus is on developing a demonstrator for real-time processing of high-rate data streams from sPHENIX experiment tracking detectors. The limitations of a 15 kHz maximum trigger rate imposed by the calorimeters can be negated by intelligent use of streaming technology in the tracking system. The approach efficiently identifies low momentum rare heavy flavor events in high-rate p+p collisions (3MHz), using Graph Neural Network (GNN) and High Level Synthesis for Machine Learning (hls4ml). Success at sPHENIX promises immediate benefits, minimizing resources and accelerating the heavy-flavor measurements. The approach is transferable to other fields. For the EIC, we develop a DIS-electron tagger using Artificial Intelligence - Machine Learning (AI-ML) algorithms for real-time identification, showcasing the transformative potential of AI and FPGA technologies in high-energy nuclear and particle experiments real-time data processing pipelines.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

The importance of cycle-by-cycle data in performing rapid battery technology development and validation

Lithium-ion battery (LiB) technology is playing a crucial role in transforming the predominantly fossil fuel-based transportation and stationary storage sectors to achieve a low-carbon economy. Rapid innovation in the LiB materials to electrode to cell design is happening to satisfy the performance, life, and safety metrics required by those myriads of applications. Lately, advanced analytics, such as machine-learning or artificial intelligence (ML/AI) techniques, are being used more frequently to aid in expedited LiB technology development, performance validation, and life prediction. The success of these techniques often relies on a large volume of well-defined and high-quality battery test data. On the other hand, most battery developers and research and development (R&D) communities are still following a classical approach to develop batteries, which is running calendar- and/or cycle-aging tests, performing reference performance tests (RPTs), and conducting post-mortem analyses periodically without paying attention to the wealth of data often not collected during the calendar or cycle life aging tests. This sparse data collection approach is time- and resource-intensive, requiring data capture and evaluation of months to years of RPT data to diagnose accurate battery state of performance, health, and safety. Even so, the underlying aging modes and mechanisms can be missed. If collected properly, battery test data during cycling or calendaring can be efficiently combined with ML/AI techniques to create powerful tools in the rapid diagnosis of battery state of performance, health, and safety along with insights into underlying aging modes and mechanisms. In this report, we discuss the importance of effective cycle-by-cycle (CBC) data collection with example case studies. Within a reasonable timeframe, RPT data are often inadequate in capturing many of the crucial battery aging dynamics, which often predominantly show up in CBC test data. Finally, we also show examples of ML/AI techniques that use CBC data in rapid diagnosis and projection of LiB state of health (SOH) to motivate the scientific community in collecting and using CBC data to facilitate expeditious technology development and validation.

25 ENERGY STORAGE

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS

Monitoring river flow status using low-cost wildlife camera and image segmentation artificial intelligence

Continuous measurement and monitoring of surface water coverage in non-perennial streams are essential for understanding the exchange fluxes between surface and subsurface waters under both inundated and non-inundated conditions. In this study, a wildlife camera photo-based framework was developed to monitor small stream water inundation, depth, discharge, and velocity. Two advanced machine learning models, YOLOv8 and Mask2Former, were utilized to efficiently analyze images captured by wildlife cameras. The accuracy of the framework was validated against on-site depth measurements at six sites in the Yakima River Basin, along with the gage height, discharge, and velocity data from four USGS sites. This approach facilitates long-term, continuous monitoring and quantification of river intermittency and water availability with high precision and low cost, thereby advancing river ecosystem research and management.

machine learning

Quantifying uncertainty in machine learning for nuclear binding energy

Techniques from artificial intelligence and machine learning are increasingly employed in nuclear theory; however, the uncertainties that arise from the complex parameter manifold encoded by the neural networks are often overlooked. Epistemic uncertainties arising from training the same network multiple times for an ensemble of initial weight sets offer a first insight into the confidence of machine learning predictions, but they often come with a high computational cost. Instead, we apply a single-model uncertainty quantification method called Δ-UQ that gives epistemic uncertainties with one-time training. Here, we demonstrate our approach on a two-feature model of nuclear binding energies per nucleon with proton and neutron number pairs as inputs. We show that Δ-UQ can produce reliable and self-consistent epistemic uncertainty estimates and can be used to assess the degree of confidence in predictions made with deep neural networks.

Huang, Mengyao [Lawrence Livermore National Labora

Data Sharing as a Catalyst for Expanding the Energy Frontier

As the energy landscape evolves to include technologies such as geothermal energy, comprehensive data become essential for driving innovation and scalability, particularly with the growing use of tools like machine learning and artificial intelligence. In emerging sectors, the cost of gathering high-quality data across large spatial areas can present a significant barrier. A key solution is leveraging existing data from well-established industries like oil and gas. However, the proprietary nature of data in these industries often hinders collaboration. This paper explores how cultivating a culture of data sharing can act as a catalyst for progress, fueling breakthroughs across both conventional and renewable energy sectors. Practical compromises that protect business interests while enabling data access are proposed, and real-world success stories are highlighted, demonstrating how collaboration has accelerated advancements in geothermal, carbon capture, and other innovative technologies.

15 GEOTHERMAL ENERGY

Optimizing enzymes for plastic upcycling using machine learning design and high throughput experiments

Plastic use is ubiquitous in the modern world, and polyethylene terephthalate (PET) is one of the most abundantly produced plastics (and the most highly produced polyester), with ~65 million metric tons manufactured annually. To the consumer, PET is likely most recognizable as the plastic used to make beverage bottles. Like many plastics, traditional mechanical or chemical means of PET deconstruction and upcycling are costly and inefficient. Because of these challenges, recycled plastic is generally of lower quality and is more expensive to produce than virgin plastic derived from petroleum. Ultimately, this results in most plastic ending up as waste. We view plastic waste as an underutilized resource which, with the development of more efficient and high-quality recycling processes, could (1) generate significant economic value while (2) decreasing petroleum usage and greenhouse gas emissions, as well as (3) minimizing its negative environmental and health impacts. Biocatalytic recycling, or biomanufacturing the basic building blocks of new plastic from plastic waste, is a promising approach to plastic reuse that complements existing recycling technologies. Recently, biological enzymes capable of breaking down PET have garnered significant attention as an attractive means of dealing with the plastic problem. These enzymes are currently undergoing pilot studies for implementation in industrial-scale enzyme-based recycling. However, there are significant limitations to current enzymes, including the need to perform costly pre-processing of the plastic waste before the enzymes are able to work. Further optimization of these enzymes is necessary to make these technologies competitive, and ultimately incentivise industry-wide adoption of this biology-based green recycling technology. n this work we demonstrate a means to design and generate performant biological enzymes, capable of efficiently deconstructing plastic waste. Specifically, we applied recent advances in artificial intelligence, machine learning, and statistical analysis to design new versions and discover natural enzymes capable of breaking down PET. We focused on optimizing key properties that are important for industrial-scale enzymatic recycling such as pH and thermotolerance. Normal testing of enzymatic plastic-deconstruction is extremely labor intensive and so through this work we also developed a robotic-assisted experimental pipeline capable of characterizing thousands of candidate enzymes. The results of this iterative, AI-guided, multi-discipline approach have led to increases in enzymatic breakdown of over 150X over starting enzymes. This work supports the rapidly developing and transformative field of biocatalytic solutions to environmental problems beyond the discovery and predictive understanding of enzymes for polymer recycling, and has wide implications for tackling numerous energy problems such as carbon capture and fixation (e.g., engineering carbon monoxide dehydrogenase and the rubisco-pathway), biomining (e.g., design of lanthanide-binding proteins) and biomanufacturing (e.g., lignin-deconstruction enzymes).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Machine learning-driven descriptions of protein dynamics at solid-liquid interfaces

This chapter has described how ML has enabled quantitative analysis of HS-AFM data to discover the physical phenomena governing protein dynamics and ordering at solid-liquid interfaces. The research detailed in this chapter modeled the rotation models of protein nanorods, the discovery of which would otherwise not be possible. By tracking the trajectories of individual protein rods from frame to frame, it was possible to model Brownian type motion and behaviors and Levy-flight dynamics that had not previously been shown. We also described the application of the Python package AtomAI, which has been developed specifically to analyze and extract physical phenomena, providing exemplar code for training an ensemble of deep neural networks to produce the semantic segmentation of AFM data and functions for encoding and decoding local environments. We last described a combinatorial approach to analyze very noisy data with a densely covered substrate where the emergence of order for the protein liquid crystals could be elucidated. By combining the methods from Case 1 and 2, it was possible to obtain the center of mass and angle for each rod in the images and track the assembly of the rods over time into a 2D liquid crystal array on the surface of mica.

protein dynamics, solid-liquid interfaces, atomic

Anion-derived contact ion pairing as a unifying principle for electrolyte design

Enabling new electrochemical technologies requires systems that can operate under ever-more demanding conditions, and progress in energy storage applications reveals tantalizing opportunities to reimagine electrolyte design for performance at extreme potentials. Here, a common thread among these innovations is the formation of significant populations of contact ion pairs (CIPs) in the electrolyte, regardless of the specific cation chemistry or solvent system. The examples summarized in this review suggest that a set of general electrolyte design rules likely exists, where the purposeful selection of anion chemistry can yield CIP structures with tunable control over reaction thermodynamics, kinetics, and interphase chemistry. Identifying the relevant descriptors for high-performance, anion-derived CIP structures can be achieved utilizing a combined experimental and computational approach, aided by machine learning and artificial intelligence, to more rapidly survey the vast combinatorial space available and to enable a new generation of electrolytes for decarbonized electrochemical processes at scale.

electrochemistry

HTESP (High-throughput electronic structure package): A package for high-throughput ab initio calculations

High-throughput ab initio calculations are the indispensable parts of data-driven discovery of new materials with desirable properties, as reflected in the establishment of several online material databases. The accumulation of extensive theoretical data through computations enables data-driven discovery by constructing machine learning and artificial intelligence models to predict novel compounds and forecast their properties. Efficient usage and extraction of data from these existing online material databases can accelerate the next stage materials discovery that targets different and more advanced properties, such as electron–phonon coupling for phonon-mediated superconductivity. However, extracting data from these databases, generating tailored input files for different ab initio calculations, performing such calculations, and analyzing new results can be demanding tasks. Here, in this work, we introduce a software package named “HTESP” (High-Throughput Electronic Structure Package) written in Python and Bash languages, which automates the entire workflow including data extraction, input file generation, calculation submission, result collection and plotting. Our HTESP will help speed up future computational materials discovery processes.

36 MATERIALS SCIENCE