Search NASA⌕ Search

SEARCH · Search NASA

Results for “AI for Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Progress in the development of the community Particle Accelerator Lattice Standard (PALS)

The Particle Accelerator Lattice Standard (PALS) is a community effort to create an open standard to promote lattice information exchange for particle accelerators. PALS development is a community-wide international effort involving accelerator physicists from multiple institutions. While it started as a lattice standard for beam dynamics simulations, it is now being extended to support other particle accelerator activities, in particular accelerator operation. With new accelerators that are becoming more complex, larger collaborations and the increasing imprint of artificial intelligence in all accelerator activities (from design to operation to workforce development), the imperative for a common, standardized accelerator ontology has been transitioning from “nice-to-have” to “must-have”. We will present the status of the project, its relations to other projects, including to two of the particle accelerator projects of the newly announced US DOE Genesis Mission: the Multi-Office Accelerator Team (MOAT) project and the Nuclear physics AI-Ready Accelerator Data (NARAD) project.

Brynes, A. [Science and Technology Facilities Coun↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Space life sciences perspectives for Space Station Freedom

It is now generally acknowledged that the life science discipline will be the primary beneficiary of Space Station Freedom. The unique facility will permit advances in understanding the consequences of long duration exposure to weightlessness and evaluation of the effectiveness of countermeasures. It will also provide an unprecedented opportunity for basic gravitational biology, on plants and animals as well as human subjects. The major advantages of SSF are the long duration exposure and the availability of sufficient crew to serve as subjects and operators. In order to fully benefit from the SSF, life sciences will need both sufficient crew time and communication abilities. Unlike many physical science experiments, the life science investigations are largely exploratory, and frequently bring unexpected results and opportunities for study of newly discovered phenomena. They are typically crew-time intensive, and require a high degree of specialized training to be able to react in real time to various unexpected problems or potentially exciting findings. Because of the long duration tours and the large number of experiments, it will be more difficult than with Spacelab to maintain astronaut proficiency on all experiments. This places more of a burden on adequate communication and data links to the ground, and suggests the use of AI expert system technology to assist in astronaut management of the experiment. Typical life science experiments, including those flown on Spacelab Life Sciences 1, will be described from the point of view of the demands on the astronaut. A new expert system, 'PI in a Box,' will be introduced for SLS-2, and its applicability to other SSF experiments discussed. (This paper consists on an abstract and ten viewgraphs.)

Young, Laurence R.↗

Enabling Interoperability in Earth System Digital Twins (ESDT): Integrating Observations, Models, and AI for Actionable Insights Through NASA'S Intelligent Systems Technology Program

NASA’s Intelligent Systems Technology Program (IST) is driving a paradigm shift in Earth science through the development of Earth System Digital Twins (ESDT). These integrated information systems create a dynamic "digital replica" of the Earth by harmonizing continuous, multi-source observations with high-fidelity models and state-of-the-art artificial intelligence (AI) that enable “What now?”, “What next?”, and “What if?” scenario building. These scenarios are reflected in NASA IST’s series of ESDTs, from the Coastal Zone Digital Twin that integrates complex data on the current state of the Chesapeake Bay to the Terrestrial Environmental Rapid-Replication and Assimilation Hydrometeorological (TerraHydro) AI-based ESDT that forecasts water movement across Earth’s surface, to the Agriculture Land Information System (AgLIS) which can be used to assess optimal planting dates and crop yield estimates. By bridging the gap between vast data archives and actionable insights, these projects enable a system-of-systems approach to understanding complex, interacting Earth processes. This poster will highlight recent innovations and future directions from NASA’s ESDT initiatives: Continuous Data Assimilation & Multi-Source Fusion. A core requirement of the ESDT work is the transition from static models to dynamic "living" replicas. This involves creating frameworks for the continual assimilation of near-real-time data from uncoordinated, heterogeneous sources, including satellite observations and airborne assets, and ground-based Internet of Things (IoT) sensors. These systems link design, operational status, and environmental data, ensuring the digital twin accurately reflects the current state of the physical Earth system. High-Fidelity Hybrid Modeling & Computational Acceleration to enable interactive "what-if" explorations, programs are moving beyond traditional, slow physical solvers by developing fast surrogate machine learning models and Deep Generative Models (DGMs). These hybrid approaches use neural networks to emulate complex physics, such as cloud feedback or ocean dynamics, at a fraction of the original computing cost, often leveraging advanced hardware like Graphics Processing Units (GPUs) to achieve the necessary scale. Federated Ecosystems & Interoperable Frameworks rather than building isolated tools, NASA IST is moving toward federated ESDTs and reusable analytic collaborative frameworks. This theme focuses on interoperability standards and common ontologies that allow specialized digital twins to interact and share data. This system-of-systems architecture supports multi-discipline investigations, such as analyzing how upstream watershed changes impact downstream urban flooding or how wildfire emissions affect regional air quality. By leveraging these advancements, ESDTs empower researchers and decision-makers to conduct real-time analysis and run complex hypothetical scenarios, ultimately improving our understanding of Earth’s evolving systems and informing critical real-world applications.

Earth System↗

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES↗

Using Large Language Models to help customers monitor global threat data

Large Language Models have proven adept at answering general knowledge questions. To make these generative AI tools useful to our mission customers for monitoring global threats, the data sciences team at Sandia is utilizing retrieval augmented generation (RAG) techniques to customize these models with local data. The local data we use consists of data such as research articles and patent abstracts that we've collected over the last several years using automated pipelines.

Herzer, John Andrew [Sandia National Laboratories ↗

SHARP: Spacecraft Health Automated Reasoning Prototype

The planetary spacecraft mission OPS as applied to SHARP is studied. Knowledge systems involved in this study are detailed. SHARP development task and Voyager telecom link analysis were examined. It was concluded that artificial intelligence has a proven capability to deliver useful functions in a real time space flight operations environment. SHARP has precipitated major change in acceptance of automation at JPL. The potential payoff from automation using AI is substantial. SHARP, and other AI technology is being transferred into systems in development including mission operations automation, science data systems, and infrastructure applications.

Atkinson, David J.↗

Usage of ChatGPT for Engineering Design and Analysis Tool Development

ChatGPT, a generative AI large language model, has recently captured significant attention in both the computer science community and the broader public domain. It has demonstrated a wide range of capabilities, from answering simple questions to writing fully functional computer code. This study spotlights both the capabilities and limitations of ChatGPT when addressing engineering problems. The model's capacity to generate practical engineering tools is highlighted through an example of a prompt that leads to an interactive plotting tool, enabling the examination of the fluid boundary layer around a fan blade. Subsequently, the paper also uncovers potential pitfalls in ChatGPT’s application, shown through an unsuccessful attempt to use ChatGPT to automate a process in Ansys Workbench through scripting. The research further investigates ChatGPT's proficiency in addressing inquiries and providing explanations about the functionalities of OpenMDAO, an open-source, multidisciplinary design, analysis, and optimization tool developed at NASA Glenn Research Center. Finally, an optimization methodology, developed with ChatGPT’s help, is applied to the structural optimization of a fan blade. The developed optimization method utilizes T-Blade3 for geometry generation, Ansys Mechanical for meshing and finite element analysis, and sci-kit learn’s MLPRegressor method to generate a trained neural network model of the design space. OpenMDAO is then used to find the optimal point within the design space. The outcome is a significant reduction in stress in the optimized model—less than one-fifth of the stress value in the baseline model.

Design↗

Usage of ChatGPT for Engineering Design and Analysis Tool Development

ChatGPT, a generative AI large language model, has recently captured significant attention in both the computer science community and the broader public domain. It has demonstrated a wide range of capabilities, from answering simple questions to writing fully functional computer code. This study spotlights both the capabilities and limitations of ChatGPT when addressing engineering problems. The model's capacity to generate practical engineering tools is highlighted through an example of a prompt that leads to an interactive plotting tool, enabling the examination of the fluid boundary layer around a fan blade. Subsequently, the paper also uncovers potential pitfalls in ChatGPT’s application, shown through an unsuccessful attempt to use ChatGPT to automate a process in Ansys Workbench through scripting. The research further investigates ChatGPT's proficiency in addressing inquiries and providing explanations about the functionalities of OpenMDAO, an open-source, multidisciplinary design, analysis, and optimization tool developed at NASA Glenn Research Center. Finally, an optimization methodology, developed with ChatGPT’s help, is applied to the structural optimization of a fan blade. The developed optimization method utilizes T-Blade3 for geometry generation, Ansys Mechanical for meshing and finite element analysis, and sci-kit learn’s MLPRegressor method to generate a trained neural network model of the design space. OpenMDAO is then used to find the optimal point within the design space. The outcome is a significant reduction in stress in the optimized model—less than one-fifth of the stress value in the baseline model.

Design↗

AI Applications to Physics Experiments at Jefferson Lab

We survey how AI/ML is being deployed across Jefferson Lab's experimental and accelerator programs. In EPSCI, Hydra applies computer vision to automate real-time data-quality monitoring across all four experimental halls, replacing manual inspection of hundreds to thousands of histograms per shift. AIEC (AI Experiment Controls) uses ML to stabilize drift chamber gains and is now part of standard CEBAF production running, while AI Optimized Polarization (AIOP) targets autonomous control of polarized targets and photon beam angular alignment. In CASA, cavity fault classification models identify faulted cavities and trip types from waveform data with ~85% and ~78% agreement to labeled data, respectively, and are deployed in production; a separate effort applies LLMs and hybrid search to make the CEBAF operations logbook AI-ready. QCD-focused work includes transformer- and GAN-based generative models for particle-level event simulation, with distributed GAN training scaling studies on Polaris. Additional efforts span ML-on-FPGA for the EIC and a new Data Science Department coordinating anomaly detection, uncertainty quantification, and HPC-scalable ML lab-wide. Collectively, these projects illustrate AI's growing role in improving efficiency across JLab's nuclear physics mission.

Mei, Xinxin [Thomas Jefferson National Accelerator↗

Creating Benchmark Data for Artificial Intelligence and Machine Learning Space Biology Research

To identify an appropriate AI/ML approach for a specific problem, the best practice is to measure algorithm performance through the benchmarking process. A scientific benchmark consists of an AI-ready dataset and a reference implementation on a specific scientific question. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML to create scientific benchmark datasets in three applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. Currently, there are no standardized datasets available to benchmark AI/ML algorithms in the domain of space biology. In this work, we constructed two AI/ML-ready biological datasets from experiments in space-flown mice: cellular imaging and RNA-seq. First, radiation-exposed immune cells harbor DNA damage foci that can be fluorescently marked to visualize the amount of damage following exposure to ionizing radiation. However, such large datasets are difficult to analyze visually, due to imaging inconsistencies and human bias, and classical image processing approaches can fail on imaging artifacts. AI/ML are therefore exciting alternative, providing the speed of machines and the accuracy of humans. We have made this dataset available at https://registry.opendata.aws/bps_microscopy/. Second, high-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. However, most sequencing datasets suffer from high dimensionality and low sample count. In this work, we used a generative adversarial network to synthesize a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data with sufficient space-flown and ground control mouse liver samples from NASA GeneLab. This dataset is available at https://registry.opendata.aws/bps_rnaseq/. These datasets are now fully open the Space Biology community to test their favorite AI/ML approaches.

James Casaletto↗

Workshop Summary Report on Using AI Tools to Improve the Efficiency and Outcomes of the NEPA Process: AI for Permitting Workshop at the 2025 National Association of Environmental Professionals (NAEP) Annual Conference

On April 29, 2025, the U.S. Department of Energy and Pacific Northwest National Laboratory hosted a workshop at the National Association of Environmental Professionals 2025 Conference and Training Symposium in Charleston, South Carolina, titled, “Effective and Responsible Use of Customized AI Tools to Improve the Efficiency and Outcomes of the NEPA Process.” The objectives of this workshop were to make environmental practitioners aware of the potential for using artificial intelligence in the National Environmental Policy Act process, demonstrate examples of how artificial intelligence can be integrated effectively to improve efficiency and outcomes and solicit questions and feedback from practitioners. This report summarizes the key points from all talks and case studies, as well as audience questions and feedback on the presentation topics and the broader topic of "AI in permitting". The report concludes by highlighting the key barriers and opportunities for the implementation of AI in permitting, as discussed during the workshop.

54 ENVIRONMENTAL SCIENCES↗

Shaping the Future of Self-Driving Autonomous Laboratories Workshop

The "Shaping the Future of Self-Driving Autonomous Laboratories" workshop, held in Denver on November 7-8, 2024, brought together leading experts from materials science and computing to address the growing need to revolutionize scientific research through AI-driven autonomous laboratories. The workshop identified critical challenges, including the integration of heterogeneous data, development of AI systems that understand fundamental physical principles, and comprehensive safety protocols. Key recommendations emerged around developing universal laboratory equipment interfaces, implementing automated metadata collection systems, and creating hybrid AI approaches that combine data-driven learning with scientific principles. The workshop emphasized maintaining human oversight while leveraging automation, transforming scientific education to prepare the next generation of researchers, and establishing a national consortium leveraging DOE facilities as anchors for broader collaboration with academia and industry. Participants stressed the urgency of addressing the growing disconnect between human decision-making timescales and modern instrumentation capabilities, highlighting the need for strategic automation while preserving essential human insight and oversight in the research process.

36 MATERIALS SCIENCE↗

Velocity Extraction Using Complete Time-Domain Waveform Data and Audio Machine Learning

We developed a new machine learning-based tool for extracting information from interferometry measurements: MIDWAZE (Modular Interferometry Direct Waveform AnalyZEr). This paper showcases MIDWAZE’s ability to extract an object’s velocity information from Photonic Doppler Velocimetry (PDV) data at near-human accuracy with little to no human intervention. MIDWAZE can extract velocities roughly 350 times as fast as a human analyst "rushing" to complete their extractions, with similar extraction accuracy. MIDWAZE’s most outstanding feature is that it operates directly in waveform/temporal space, freeing analysis from certain limitations imposed by traditional spectrogram-based approaches and opening the way to "phase aware" PDV analysis. MIDWAZE also has limited ability to discriminate between different solid objects, which we develop as a first step towards automated discrimination of different kinds of objects such as ejecta clouds.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES↗