Search NASASearch

SEARCH · Search NASA

Results for “Responsible AI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Responsible AI Recommendations (RAIR)

This dataset contains recommendations culled from academic and gray literature on the implementation of responsible or ethical AI. Recommendations are tagged with 1 to 5 topical labels, and a label regarding the source type of their parent document.

Artificial intelligence

Responsible Adoption of Artificial Intelligence (AI) in Electric Grid Operations

The future of the grid will be powered by AI—or undermined by it. Artificial intelligence is rapidly reshaping grid operations, improving fault detection, forecasting accuracy, and real-time optimization. As AI systems move closer to operational decision loops, however, they introduce new consequence pathways: expanded attack surfaces, model integrity risks, regulatory exposure, and human-automation challenges. This talk presents a consequence-driven framework for deploying AI responsibly in the electric grid. Attendees will gain practical strategies to strengthen resilience, boost reliability, and deploy AI securely — ensuring the grid of the future is not only smarter but safer.

25 - ENERGY STORAGE

Radio Afterglow Detection and AI-driven Response (RADAR): A Federated Framework for Gravitational-wave Event Follow-up

The landmark detection of both gravitational waves (GWs) and electromagnetic (EM) radiation from the binary neutron star merger GW170817 has spurred efforts to streamline the follow-up of GW alerts in current and future observing runs of ground-based GW detectors. Within this context, the radio band of the EM spectrum presents unique challenges. Sensitive radio facilities capable of detecting the faint radio afterglow seen in GW170817, and with sufficient angular resolution, have small fields of view compared to typical GW localization areas. Additionally, theoretical models predict that the radio emission from binary neutron star mergers can evolve over weeks to years, necessitating long-term monitoring to probe the physics of the various postmerger ejecta components. These constraints, combined with limited radio observing resources, make the development of more coordinated follow-up strategies essential—especially as the next generation of GW detectors promises a dramatic increase in detection rates. Here, we present RADAR, a framework designed to address these challenges by promoting community-driven information sharing, federated data analysis, and system resilience, while integrating AI methods for both GW signal identification and radio data aggregation. We show that it is possible to preserve data rights while sharing models that can help design and/or update follow-up strategies. We demonstrate our approach through a case study of GW170817, and discuss future directions for refinement and broader application.

Gravitational waves

Generative AI for Power Grid Operations

Generative artificial intelligence (AI) has captured into the mainstream, demonstrating capabilities that once belonged solely to the realm of human cognition. From defeating world champions in complex games to generating human-quality text and images, Generative AI has proven its potential to revolutionize countless industries. The electric power grid is no exception. Generative AI's ability to process vast amounts of data rapidly, assist decision support and identify patterns could significantly enhance power grid operations. For example, Generative AI could improve state estimation where measurements are not available or integrate renewable energy sources more efficiently with probabilistic forecasting. The key contributions of this whitepaper are outlined below: (1) Comprehensive overview of Generative AI's applications in power grid operations: It highlights the opportunities in areas such as forecasting, state estimation, and demonstrating the potential for enhancing efficiency, reliability, and resilience. (2) Expanding Generative AI's impact through synergies with emerging technologies: The paper introduce NREL developed eGridGPT and explores how AI orchestration, multi-agent systems, and Digital Twins can collaborate to optimize grid operations, addressing the complexities of a decarbonized and electrified future. (3) In-depth analysis of challenges in implementing Generative AI: This includes considerations like data availability and quality, model validation, certification, and ethical concerns, ensuring responsible AI deployment. (4) Emphasizing human-AI collaboration: The whitepaper underscores the importance of trustworthy, transparency, and explainability in AI systems to promote seamless interaction between human operators and AI, ultimately improving decision-making. (5) Exploring future research and development: It identifies critical areas for further advancement to fully realize Generative AI's potential in power grid operations. This whitepaper serves as a valuable resource for researchers, practitioners, and policymakers looking to harness Generative AI for a more reliable, stable, and cost-effective power grid.

24 POWER TRANSMISSION AND DISTRIBUTION

The Perils of AI-Assisted Time Conversion Critical Errors in J2000 Time to UNIX/UTC Conversions Using AI Assistants (Technical Memorandum)

This memo addresses a critical issue encountered when using AI assistants for time system conversions, specifically converting data tagged with J2000 seconds through UNIX time to UTC. Initial AI responses to this common conversion task were fundamentally incorrect, with errors exceeding 37 seconds - enough to cause significant issues in many applications. This document demonstrates the problem, quantifies the error, and provides best practices for using AI tools safely in technical work. Both Claude 4.5 Sonnet and ChatGPT GPT-5.5 Thinking gave the same answer.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

This report summarizes insights from the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science, which convened more than 40 experts from national laboratories, academia, industry, and community organizations to chart a path toward more powerful, sustainable, and collaborative scientific software ecosystems. To address urgent challenges at the intersection of high-performance computing (HPC), AI, and scientific software, participants envisioned agile, robust ecosystems built through socio-technical co-design—the intentional integration of social and technical components as interdependent parts of a unified strategy. This approach combines advances in AI, HPC, and software with new models for cross-disciplinary collaboration, training, and workforce development. Key recommendations include building modular, trustworthy AI-enabled scientific software systems; enabling scientific teams to integrate AI systems into their workflows while preserving human creativity, trust, and scientific rigor; and creating innovative training pipelines that keep pace with rapid technological change. Pilot projects were identified as near-term catalysts, with initial priorities focused on hybrid AI/HPC infrastructure, cross-disciplinary collaboration and pedagogy, responsible AI guidelines, and prototyping of public-private partnerships. This report presents a vision of next-generation ecosystems for scientific computing where AI, software, hardware, and human expertise are interwoven to drive discovery, expand access, strengthen the workforce, and accelerate scientific progress.

97 MATHEMATICS AND COMPUTING

Poster: Responsible Adoption of Artificial Intelligence (AI) in Electric Grid Operations

The rapid integration of artificial intelligence (AI) in the utility transmission and distribution (T&D) sector is revolutionizing traditional grid management practices. As utilities encounter complexities from evolving consumer behaviors and energy integration, AI becomes a critical solution for enhancing grid monitoring, fault detection, and operational optimization. However, increased reliance on interconnected technologies introduces significant cybersecurity risks, regulatory compliance challenges, and human factors concerns. This study proposes a strategic, responsible and consequence-driven approach to AI implementation, examining the dual nature of AI adoption by highlighting its transformative benefits for utilities and associated risks. It provides utilities with a framework for evaluating AI integration, enabling them to navigate challenges and capitalize on opportunities to achieve greater reliability, efficiency, and resilience in an increasingly complex energy landscape.

24 - POWER TRANSMISSION AND DISTRIBUTION

Agentic framework for programmatic crystal structure generation using a fine-tuned worker–supervisor large language model

Platinum group metals (PGMs) underpin many catalytic technologies but face severe supply constraints, motivating the search for alternative materials and computational methods to accelerate discovery. While atomistic simulation tools such as Pymatgen and ASE have streamlined structure manipulation, they require detailed inputs, limiting accessibility for experimentalists and slowing early-stage exploration. Here, in this study, we present an AI-driven agentic framework that orchestrates worker–supervisor large language models (LLMs). The worker translates natural-language prompts of varying abstraction into valid crystallographic structures using a compact LLM fine-tuned with low-rank adaptation on a curated text–code–CIF dataset, emphasizing energy-efficient training. Benchmarking against the baseline CodeGen-350M-mono model shows that fine-tuning reduces hallucination rates from 100% to as low as 5% and improves structural match accuracy to up to 82% for fully specified inputs. Accuracy declines with decreasing prompt detail but remains nontrivial even when only stoichiometry and space group are provided, underscoring the LLM’s capacity for crystallographic inference. The supervisor Claude LLM evaluates the outputs and triggers iterative refinement through the worker’s built-in structure manipulation capabilities (e.g., supercell scaling, strain, vacancy, and substitution operations). We further demonstrate use cases for technologically relevant catalysts, including IrO 2 , pyrochlore Pb 2 Ir 2 O 7 , Ni 2 FeO 4 , and Ni 3 Mo, where the framework generates physically consistent structures that can be refined via geometry optimization. This work introduces a low-energy, language-driven pathway for integrating human and machine intelligence in materials design, paving the way for AI-assisted synthesis planning and high-throughput screening of complex oxides.

AI agent

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]

ADEPT: A Pedagogical Framework for Integrating Agentic AI with Deterministic Scientific Workflows

The integration of Large Language Models (LLMs) into scientific research promises to accelerate discovery, yet a significant gap remains between the dynamic reasoning of Artificial Intelligence (AI) agents and the static, deterministic nature of canonical scientific workflows. This paper introduces ADEPT (Agentic Discovery and Exploration Platform for Tools), a reference architecture and pedagogical framework explicitly designed to bridge this gap. ADEPT's primary mission is to provide a transparent, "glass-box" environment where researchers and engineers can learn to effectively wrap established scientific software (e.g., BLAST, Nextflow pipelines) and compose it into reliable, agent-driven workflows. We describe its modular, multi-server architecture, which leverages the Model Context Protocol (MCP) for tool serving, LangGraph for robust agentic orchestration, and a secure nsjail-based sandbox for safe code execution. By prioritizing architectural clarity, safety, and modularity, ADEPT serves as an extensible blueprint for building trustworthy AI-augmented systems and fosters the collaborative development necessary to responsibly employ agentic AI for science. We provide practical examples of how to adapt and extend this framework, highlighting its utility in workforce development and AI-readiness capabilities across research and development projects.

97 MATHEMATICS AND COMPUTING

Workshop Summary Report on Using AI Tools to Improve the Efficiency and Outcomes of the NEPA Process: AI for Permitting Workshop at the 2025 National Association of Environmental Professionals (NAEP) Annual Conference

On April 29, 2025, the U.S. Department of Energy and Pacific Northwest National Laboratory hosted a workshop at the National Association of Environmental Professionals 2025 Conference and Training Symposium in Charleston, South Carolina, titled, “Effective and Responsible Use of Customized AI Tools to Improve the Efficiency and Outcomes of the NEPA Process.” The objectives of this workshop were to make environmental practitioners aware of the potential for using artificial intelligence in the National Environmental Policy Act process, demonstrate examples of how artificial intelligence can be integrated effectively to improve efficiency and outcomes and solicit questions and feedback from practitioners. This report summarizes the key points from all talks and case studies, as well as audience questions and feedback on the presentation topics and the broader topic of "AI in permitting". The report concludes by highlighting the key barriers and opportunities for the implementation of AI in permitting, as discussed during the workshop.

54 ENVIRONMENTAL SCIENCES

From Rules to Reasoning: A Survey of Large Language Model-Based Approaches to Scientific Hypothesis and Idea Generation

Scientific hypothesis generation represents a fundamental challenge in contemporary research due to exponentially expanding literature volumes and increasing disciplinary specialization. Large language models (LLMs) have emerged as transformative tools for automated scientific discovery, moving beyond traditional rule-based and literature-mining approaches. Four paradigmatic approaches define current LLM-driven hypothesis generation: direct prompting and fine-tuning methods, knowledge-enhanced frameworks integrating retrieval-augmented generation (RAG), multi-agent collaborative systems simulating research teams, and reasoning-focused approaches implementing cognitive architectures. Domain-specific applications demonstrate statistical equivalence to human expert performance in social psychology, experimental validation in biomedical research, and near-expert quality in astronomy. Evaluation methodologies encompass human expert assessment, LLM-as-judge frameworks, and comprehensive benchmarking systems. Technical challenges include hallucination management, knowledge integration limitations, and balancing novelty with feasibility. Future directions emphasize hybrid neural-symbolic architectures and sophisticated human-AI collaboration models for responsible scientific discovery acceleration.

AI-driven discovery

Yes, No, Maybe So: Human Factors Considerations for Fostering Calibrated Trust in Foundation Models Under Uncertainty

High-stakes analytical environments require analysts to evaluate evidence and generate conclusions to inform critical decisions often under conditions of uncertainty. Probabilistic decision-making based on incomplete or inaccurate information can reduce productivity, compromise national interests, and endanger public safety. Researchers are developing expert systems built on foundation models (FMs) to support analysts’ decision-making processes by enabling human-artificial intelligence (AI) teaming, in part through the quantification and expression of uncertainty information. As FMs continue to mature, it is imperative to correspondingly consider analysts’ needs for appropriately interpreting and using uncertainty information. However, prior research indicates that it remains unclear how analysts engage with FM-generated uncertainty information and the extent to which these interactions influence trust in, and reliance on, expert systems. We plan to review the state of the science and conduct an exploratory, qualitative study to (a) understand how properly communicated uncertainty can foster calibrated trust and appropriate reliance and (b) identify approaches for effectively conveying FM-generated uncertainty information during analytical workflows. We will administer semi-structured interviews with analysts from a specific high-stakes analytical environment to collect their current experiences with job-related uncertainty and their impressions when viewing FM-generated uncertainty information. During the interview protocol, participants will be presented with several different FM outputs and invited to discuss their thoughts and beliefs about the uncertainty information displayed. Participants may provide insights into how trust and reliance may be influenced by uncertainty. The results of this study will help us to better understand how analysts currently interpret and use uncertainty information. Our findings may inform human factors recommendations for effectively conveying uncertainty information to foster calibrated trust in, and appropriate reliance on, expert systems. Interaction designers and FM developers can use this knowledge to enhance human-AI teaming and ensure the responsible deployment of FM-based expert systems in analytical workflows.

97 MATHEMATICS AND COMPUTING

Agnostic capture of pathogens for the detection and diagnostics of emerging threats

The continued emergence of pathogens, whether novel, re-emerging, or engineered, poses a persistent global biosecurity and public health challenge. Recent outbreaks, including COVID-19, Lassa fever, Marburg virus, mpox, and avian influenza, underscore the urgent need for robust systems that enable rapid surveillance, early diagnosis, and timely countermeasures before widespread human transmission occurs. In this article, we focus on early detection technologies and systematically evaluate current diagnostic and sensing modalities. We highlight sequencing and spectroscopy as two complementary approaches capable of providing broad, agnostic detection and rich biological insight. Our analysis emphasizes that scientific innovation alone is insufficient: effective preparedness also requires improved data curation, integration, and sharing to build AI-ready resources that accelerate future responses. We argue for coordinated advances in both technological capabilities and supporting infrastructure to enable the rapid identification and characterization of emerging pathogens and to fully leverage modern science against evolving infectious threats.

Environmental health

Modernizing Accelerator Responsiveness and Controls in Operations

Accelerators increasingly use artificial intelligence (AI) and machine learning (ML) software and workflows for a variety of tasks, from optimization to fault detection and recovery. Efficient and sustainable application of these technologies necessitates specialized and facility-specific infrastructure commitments. Accelerator facilities also introduce unique radiation and security hazards, placing additional demands on operational infrastructure. These needs further escalate the prioritization of effective collaboration models and associated funding mechanisms and legal frameworks.

43 PARTICLE ACCELERATORS

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING