Search NASA⌕ Search

SEARCH · Search NASA

Results for “generative artificial intelligence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Discovery and problem solving: Triangulation as a weak heuristic

Recently the artificial intelligence community has turned its attention to the process of discovery and found that the history of science is a fertile source for what Darden has called compiled hindsight. Such hindsight generates weak heuristics for discovery that do not guarantee that discoveries will be made but do have proven worth in leading to discoveries. Triangulation is one such heuristic that is grounded in historical hindsight. This heuristic is explored within the general framework of the BACON, GLAUBER, STAHL, DALTON, and SUTTON programs. In triangulation different bases of information are compared in an effort to identify gaps between the bases. Thus, assuming that the bases of information are relevantly related, the gaps that are identified should be good locations for discovery and robust analysis.

Rochowiak, Daniel↗

Combining Generative Modeling and Advanced Control for Building Scenario Generation

Buildings make up a large portion of energy consumption in the U.S. today. Understanding their energy consumption patterns can improve their efficiency, but requires detailed models that rely on incomplete or unknown information. Previous work has shown that artificial intelligence (AI) can be used to predict missing information and even suggest upgrades to improve building efficiency. However, building upgrades may require undesirable upfront costs. Oppositely, advanced control could improve building efficiency with negligible upfront cost. To explore the tradeoffs between these two approaches, in this work we propose a workflow to compute optimal temperature setpoint schedules to minimize energy consumption and operational cost. Results show that modifying the temperature setpoints in a building using model predictive control (MPC) can effectively reduce its energy consumption and operational cost. This optimal operation cannot fully meet a desired goal. However, we show that by considering MPC in addition to component upgrades, a desired goal can be met with significantly less upfront costs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The role of AI in detecting and mitigating human errors in safety-critical industries: A review

For safety-critical industries, human error (HE) presents continual risks to system productivity, reliability and safety. Artificial intelligence (AI) and machine learning (ML) methods have emerged as promising approaches to understand, categorize and mitigate the risk of HE in safety-critical industries. Furthermore, this review offers an examination of the current landscape regarding the utilization of AI/ML with regards to HE in safety-critical industries, categorizing literature into descriptive modeling, predictive modeling, prescriptive modeling, and generative modeling techniques. Additionally, the review aims to provide insights regarding themes in literature, challenges, and future research directions. Findings of the review suggest that AI/ML methods can prove useful in addressing the HE problem across safety-critical industries.

42 ENGINEERING↗

An intelligent interface for satellite operations: Your Orbit Determination Assistant (YODA)

An intelligent interface is often characterized by the ability to adapt evaluation criteria as the environment and user goals change. Some factors that impact these adaptations are redefinition of task goals and, hence, user requirements; time criticality; and system status. To implement adaptations affected by these factors, a new set of capabilities must be incorporated into the human-computer interface design. These capabilities include: (1) dynamic update and removal of control states based on user inputs, (2) generation and removal of logical dependencies as change occurs, (3) uniform and smooth interfacing to numerous processes, databases, and expert systems, and (4) unobtrusive on-line assistance to users of concepts were applied and incorporated into a human-computer interface using artificial intelligence techniques to create a prototype expert system, Your Orbit Determination Assistant (YODA). YODA is a smart interface that supports, in real teime, orbit analysts who must determine the location of a satellite during the station acquisition phase of a mission. Also described is the integration of four knowledge sources required to support the orbit determination assistant: orbital mechanics, spacecraft specifications, characteristics of the mission support software, and orbit analyst experience. This initial effort is continuing with expansion of YODA's capabilities, including evaluation of results of the orbit determination task.

Schur, Anne↗

Intern Poster

Large Language Models (LLMs) have skyrocketed in popularity after the release of ChatGPT in late 2022. Although LLMs are powerful tools, they can be subject to hallucinations, which is when an LLM (or any AI model) produces misleading/ nonsensical information. The objective is to determine if statistical methods can be used to detect hallucinations as an LLM generates its answer token by token (essentially word by word).

97 - MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields, in part due to a culture of open data sharing and reuse. AI/ML methodology is well-suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Inexperienced researchers can produce models that perform poorly outside of the training dataset. Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Casaletto↗

Artificial intelligence tools for enzyme engineering and metabolic engineering

Enzyme engineering and metabolic engineering drive innovation in energy biotechnology. In recent years, artificial intelligence (AI) has supported successful applications in designing effective enzymes and productive microbial cell factories. This review summarizes recent advances in enzyme redesign using protein language models, de novo enzyme design with generative models, and AI tools for engineering metabolism and related cellular phenotypes. Across these areas, AI models are shifting from single modality inputs to integrated representations of protein function, metabolic pathways, and cell states. We emphasize that unifying the diverse data representations across scales will be necessary for advancements in energy biotechnology.

Volk, Michael [Univ. of Illinois at Urbana-Champai↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

A roadmap toward scaling, reasoning and self-evolving foundation models for nuclear and particle physics

Foundation models have revolutionized artificial intelligence, with Large Language Models demonstrating unprecedented capabilities in multimodal understanding, reasoning and tool use. Nuclear and particle physics stands at a critical juncture where similar transformative potential awaits realization. The field generates exabytes of experimental data, exascale simulations, and decades of theoretical insights — yet these remain largely disconnected from modern Artifical Intelligence (AI) capabilities, with most physics AI applications confined to narrow, task-specific models that suffer from domain shifting when applied to real experimental data. We present a roadmap for FM4NPP (Foundation Model for Nuclear and Particle Physics), systematically scaling from current proof-of-concept models to trillion-parameter architectures capable of autonomous discovery. Our approach advances three critical frontiers: unified data infrastructure integrating detector data, scientific knowledge and computational tools across global facilities; multi-facility foundation models enabling cross-experiment knowledge transfer and accelerated discovery; and agentic AI capabilities for reasoning and autonomous tool use. The resulting self-evolving FM4NPP will transform physics research by converting time-intensive data analysis, theory derivation and computational bottlenecks into rapid AI–human collaborative discovery. This paradigm shift promises to fundamentally accelerate scientific progress in nuclear and particle physics, enabling researchers to focus on high-level insights while AI handles routine analysis and explores vast parameter spaces beyond human capacity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Solar PV

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) research platform. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence data centers and other variable loads. This dataset entry describes the behavior of a 1.25-MW proton exchange membrane MC250 electrolyzer system, manufactured by Nel Hydrogen , [1] when fed historical data generated by the 430-kW, fixed-axis solar photovoltaic (PV) array located at NLR’s Flatirons Campus. (While the electrolyzer balance of plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack.) Solar PV power output data for the 2020 calendar year were categorized on a daily basis by total energy generation and standard deviation. Each day was then ranked by these metrics, and the 25th, 50th, and 100th percentiles were selected. The 75th percentile day did not exhibit sufficient variability to make for a valuable experiment. A similar process was used for the related historical wind dataset . [2] The historical days in 2020 that represented these percentiles are Dec. 19, March 29, and May 4, respectively. The entire solar day’s power profile was then fed through the MC250 electrolyzer. Due to its length, the 100th percentile day experiment was split into two parts, and the final 3 hours of the solar day were not captured. These final 3 hours contained no spikes or dips of interest and simply represented a slow decay of input solar power. Also, a single timestamp (13:13:47 on Jan. 14, 2026) was lost in the hydrogen system supervisory control and data acquisition. Finally, during the 25th percentile experiment (solar day Dec. 19, 2020) data recording was lost from 11:00:13 to 11:14:45. The roughly 15 minutes of the solar profile were rerun at the end of the experiment and spliced into this time slot during post-processing. The electrolysis system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operation of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical solar profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. For more details on the statistical analysis process, see the slide deck “Public Reference Data for Megawatt-Scale Hydrogen Electrolysis: NLR Historical Solar PV Analysis and Profile Generation” accessible with this data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single solar PV electrolysis experiment and is formatted as: {technology}_{percentile}_{scaling factor} For instance, “solarPV-430kW_25_2x.zip” reports the experiment using the 25th percentile solar data from the historical 2020 solar PV dataset, scaled to 200%. Scaling factors were applied to the generated solar PV power output files to more closely match the 1.25-MW capacity of the electrolyzer. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and solar power input. A PDF file detailing the historical solar data statistical analysis used to generate the solar profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all experiments combined into one dataset labeled "combined_solarPV_experiments.csv". [1] nelhydrogen.com/product/mc-series-electrolyser . [2] data.nlr.gov/submissions/316 .

08 HYDROGEN↗

Artificial Intelligence Transforming Post-Translational Modification Research

Post-Translational Modifications (PTMs) are covalent changes to amino acids that occur after protein synthesis, including covalent modifications on side chains and peptide backbones. Many PTMs profoundly impact cellular and molecular functions and structures, and their significance extends to evolutionary studies as well. In light of these implications, we have explored how artificial intelligence (AI) can be utilized in researching PTMs. Initially, rationales for adopting AI and its advantages in understanding the functions of PTMs are discussed. Then, various deep learning architectures and programs, including recent applications of language models, for predicting PTM sites on proteins and the regulatory functions of these PTMs are compared. Finally, our high-throughput PTM-data-generation pipeline, which formats data suitably for AI training and predictions is described. We hope this review illuminates areas where future AI models on PTMs can be improved, thereby contributing to the field of PTM bioengineering.

59 BASIC BIOLOGICAL SCIENCES↗

SUBTASK 1.6 – BASIN ELECTRIC CARBON STORAGE RESEARCH PROJECT: NOVEL MONITORING TECHNIQUES

The Energy & Environmental Research Center (EERC) conducted baseline activities associated with an applied research project at Basin Electric Power Cooperative’s (Basin’s) carbon capture and storage (CCS) site in Beulah, North Dakota, to establish novel carbon storage-monitoring techniques as commercial methods under Cooperative Agreement No. DE-FE0024233, Subtask 1.6. The following report summarizes the baseline activities performed and briefly describes the subsequent (operational monitoring) activities that have been proposed to the U.S. Department of Energy (DOE) as part of the overall project to develop and demonstrate novel monitoring techniques at North America’s largest permitted CCS operation. Dakota Gasification Company (DGC), a wholly owned subsidiary of Basin, owns and operates the Great Plains Synfuels Plant (GPSP) approximately 5 miles northwest of the town of Beulah, North Dakota (Figure 1). In 2023, DGC received approval from the North Dakota Industrial Commission (NDIC) to develop a storage facility on-site for injecting a stream of carbon dioxide (CO2) captured from GPSP. DGC will transport the captured CO2 stream with approximately 6.8 miles of transmission lines that extend north of GPSP and inject >1 million tonnes (MMt) of CO2 annually (>1 MMt/yr) over a 12-year period with up to six underground injection control (UIC) Class VI-compliant injection wells completed in the Broom Creek Formation, a predominantly sandstone reservoir and saline aquifer underlying GPSP. The Broom Creek Formation lies approximately 5900 feet (ft) below ground surface (bgs) at GPSP. The commercial scale (i.e., >1 MMt/yr) of DGC’s permitted carbon storage project is ideal for developing and testing the novel monitoring techniques included within Subtask 1.6. The goals of this project are to demonstrate 1) the cost-effectiveness of novel monitoring technologies included as part of this research, 2) technology capability for tracking the CO2 plume and/or associated pressure response in the subsurface and monitoring out-of-zone migration, and 3) compliance with UIC Class VI program requirements. The research activities proposed for the overall project include 1) design of an automated, integrated, modular (AIM) monitoring station; 2) time-lapse electromagnetic (EM) field surveys; 3) drone-based surveillance studies; 4) time-lapse monitoring with seismic methods; 5) advanced wellbore-monitoring methods; 6) deployment of an AIM monitoring network; 7) EM monitoring of CO2 with real-time data processing; 8) continued seasonal drone-based surveillance studies; 9) seismic monitoring with passive and active surveys; and 10) wellbore monitoring with nuclear magnetic resonance (NMR) for near-surface characterization. Completion of Activities 1.0–5.0 (baseline activities) are described in this report. Upon authorization of funding by DOE, the EERC will initiate Activities 6.0– 10.0 (operational monitoring activities). Current state-of-the-art (SOA) carbon storage-monitoring techniques require countless labor hours dedicated to the acquisition of data. Once data are gathered, these SOA techniques often rely on commercial facilities to process raw data from the field. However, it is anticipated that next-generation monitoring techniques, such as those being demonstrated, will lower acquisition footprints, be less operationally intensive, and improve data acquisition efficiencies. These new techniques are more conducive to the application of machine learning, artificial intelligence, and automation, thus providing a pathway for integration into active control systems, informing site operability, and improving the integration of data for future CCS projects across the United States. Additionally, reclaimed and active mining lands are present within the project site, creating a unique opportunity to demonstrate the effectiveness of remote sensing and surface-based geophysics monitoring techniques at similar project sites that may include disturbed, unconsolidated, or actively excavated near-surface environments. The efforts included in the overall project will produce necessary designs, learnings, and data acquired during the baseline and operational monitoring periods that are necessary for time-lapse demonstration and validation of the described monitoring techniques. In addition, it is anticipated that the monitoring technologies included in this study will be compliant with UIC Class VI requirements to enable the potential for implementation at other CCS sites across the United States.

42 ENGINEERING↗

Revealing the Hidden Third Dimension of Point Defects in Two-Dimensional MXenes

Point defects govern many important functional properties of two-dimensional (2D) materials. However, resolving the three-dimensional (3D) arrangement of these defects in multi-layer 2D materials remains a fundamental challenge, hindering rational defect engineering. Here, we overcome this limitation using an artificial intelligence-guided electron microscopy workflow to map the 3D topology and clustering of atomic vacancies in Ti3C2TX MXene. Our approach reconstructs the 3D coordinates of vacancies across hundreds of thousands of lattice sites, generating robust statistical insight into their distribution that can be correlated with specific synthesis pathways. This large-scale data enables us to classify a hierarchy of defect structures-from isolated vacancies to nanopores-revealing their preferred formation and interaction mechanisms, as corroborated by molecular dynamics simulations. This work provides a generalizable framework for understanding and ultimately controlling point defects across large volumes, paving the way for the rational design of defect-engineered functional 2D materials.

2D materials↗

Adding intelligence to scientific data management

NASA plans to solve some of the problems of handling large-scale scientific data bases by turning to artificial intelligence (AI) are discussed. The growth of the information glut and the ways that AI can help alleviate the resulting problems are reviewed. The employment of the Intelligent User Interface prototype, where the user will generate his own natural language query with the assistance of the system, is examined. Spatial data management, scientific data visualization, and data fusion are discussed.

Campbell, William J.↗

Next Generation User Support Tools

One of the manually intensive efforts of Hubble Space Telescope (HST) observing is the specification and validation of the detailed proposals for scientists observing with the telescope. In order to meet the operational cost objectives for the Next Generation Space Telescope (NGST), this process needs to be dramatically less time consuming and less costly. We have prototyped a new proposal development system, the Scientist's Expert Assistant (SEA), using a combination of artificial intelligence and user interface techniques to reduce the time and effort involved for both scientists and the telescope operations staff. The Advanced Architectures and Automation Branch of NASA's Goddard Space Flight Center is working with the Space Telescope Science Institute (ST ScI) to explore SEA alternatives. We are testing the usefulness of rule-based expert systems to painlessly guide a scientist to his or her desired observation specification. We are also examining several potential user interface paradigms and exploring data visualization schemes to see which techniques are more intuitive. Our prototypes will be validated using HST's Advanced Camera for Surveys (ACS) instrument (scheduled for installation in 1999) as a live test instrument. Having an operational test-bed will ensure the most realistic feedback possible for the prototyping cycle. In addition, once the NGST instruments are better defined, the SEA will already be a proven platform that simply needs adapting to NGST-specific instruments.

Jones, Jeremy E.↗