Search NASASearch

SEARCH · Search NASA

Results for “Agent”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

MCP-enabled agentic AI workflow for building energy modelling: framework and use cases

Traditional building energy modelling workflows remain labor-intensive and error-prone, requiring specialized expertise that limits broader adoption. This paper introduces a novel Model Context Protocol (MCP)-enabled framework that connects AI assistants to EnergyPlus through MCP, a standardized interface for tool invocation and context management. Two complementary integration paradigms are presented and compared: conversational integration, where users interact through natural language while an AI assistant orchestrates MCP tools on demand, and agentic workflow integration, where specialized agents coordinate autonomously to complete multi-step tasks. Using an experimental testbed for residential buildings, the end-to-end workflows are demonstrated. The conversational approach reduced typical inspection and modification tasks from 1-2 h to under 15 min, while maintaining full transparency through visible tool invocations. The agentic approach automated parametric analysis. These demonstrations establish MCP as a foundational layer for AI-assisted building energy modelling, enabling natural language interactions with simulation tools while preserving professional oversight and decision-making authority.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Multi-agent AI collaboration for digital twin development and assessment

Developing a digital twin (DT) model involves different steps that encompass formulating requirements, model development, implementation, and assessment with respect to real applications. Human expertise is required to coordinate and implement different steps in the DT development and assessment process. However, certain parts of this process can be automated using artificial intelligence (AI) agents for efficient workflow development. In this work, we test and analyze a multiagent AI collaboration with humans in the loop to automate different elements of the DT development and assessment process. To implement the workflow for multiagent AI DT development and assessment, we use Autogen, a multiagent framework developed by Microsoft. Autogen offers a modular and flexible framework for configuring and designing task-specific multiagent workflows. In this framework, large language models (LLMs) form the core intelligence of the AI agents where the quality and performance of the automated element is governed by the inherent capabilities and knowledge base of the LLM. We use retrieval augmented generation to supplement the LLM with relevant domain-specific information for DT requirement formulation. We illustrate this multiagent workflow using a case study on a thermal energy storage system, focusing on how AI agents can collaborate with humans to expedite and optimize different elements of DT development and assessment process.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Evaluation of dried blood spot sampling for verification of exposure to chemical threat agents

Abstract Purpose Exposure to chemical threat agents (CTAs), including nerve agents, the vesicating agent sulfur mustard, and opioids, remains a significant threat to warfighter and civilian populations. Definitive analytical methods to verify exposure to CTAs require shipping refrigerated or frozen biomedical samples to reference laboratories for analysis. Logistical and financial burdens arise as the transport of biomedical samples is subject to strict restrictions and complex packaging, which, if done incorrectly, can lead to sample deterioration. The use of dried blood spot (DBS) sampling could provide operational improvements for collecting, storing, and shipping important forensic samples. Therefore, this effort focuses on developing DBS techniques with Mitra® 30-µL volumetric absorptive microsampling (VAMS®) devices for use in CTA exposure verification. Methods VAMS® devices were loaded and dried with human whole blood that was exposed to the metabolites pinacolyl methylphosphonic acid (PMPA), ethyl methylphosphonic acid (EMPA), 1,1’sulfonylbis[2-(methylsulfinyl)ethane] (SBMSE), norfentanyl, norcarfentanil, norsufentanil, and norlofentanil. Following extraction from the VAMS® devices, metabolites were detected using liquid chromatography-tandem mass spectrometry (LC–MS/MS). The methods were validated for performance by assessing sensitivity, precision, accuracy, and recovery. Results These methods were sensitive to 1 ng/mL for SBMSE, 0.5 ng/mL for PMPA, EMPA, and norfentanyl; 0.1 ng/mL for norlofentanil, and 0.05 ng/mL for norsufentanil and norcarfentanil. All methods met acceptable precision and accuracy criteria with favorable recovery. Conclusions These results demonstrated the utility of VAMS® in stabilizing human whole blood and show promise as an improved collection method for verification of exposure to various CTAs.

Toxicology

ChemGraph as an agentic framework for computational chemistry workflows

Atomistic simulations are essential in chemistry and materials science but remain challenging to run due to the expert knowledge required for the setup, execution, and validation stages of these calculations. We present ChemGraph, an agentic framework powered by artificial intelligence and state-of-the-art simulation tools to streamline and automate computational chemistry and materials science workflows. ChemGraph leverages graph neural network-based foundation models for accurate yet computationally efficient calculations and large language models (LLMs) for natural language understanding, task planning, and scientific reasoning to provide an intuitive and interactive interface. We evaluate ChemGraph across 13 benchmark tasks and demonstrate that smaller LLMs (GPT-4o-mini, Claude-3.5-haiku, Qwen-2.5-14B) perform well on simple workflows, while more complex tasks benefit from using larger models. Importantly, we show that decomposing complex tasks into smaller subtasks through a multi-agent framework enables GPT-4o to reach perfect accuracy and smaller LLMs to match or exceed single-agent GPT-4o's performance in these benchmarks.

Computational chemistry

Deep Multi-Agent Reinforcement Learning for Real-World Signalized Traffic Corridor Control

Signalized traffic control problem has been addressed recently with deep Reinforcement Learning (RL) approaches involving diverse state, action, and reward structures. While significant progress has been noted in the literature, open challenges still remain in the areas of adaptive signal phase timing, coordination in a multi-intersection corridor setting, and consideration of real-world traffic conditions. In the context of deep RL-based problem framing, extensions are needed that enable adaptive signal phase timings in an intersection agent's action space, computationally efficient information sharing among neighboring signalized intersection agents along a corridor, and experimentation in realistic simulation environments. In this paper, we develop a deep Advantage Actor Critic (A2C) multi-agent RL (MARL) approach capturing the research extensions above and apply it within a real-world calibrated Aimsun Next traffic corridor simulation model based on traffic data from the City of Coral Gables, Florida. For a multi-intersection corridor control setting, our numerical simulation experiments with a decentralized A2C MARL algorithm applied at different time periods led to a total average corridor travel delay reduction (expressed in seconds/mile averaged over vehicles) from 4.9% to 19.9% compared to state-of-the-art actuated control.

Shuvo, Salman S. [BATTELLE (PACIFIC NW LAB)]

Three‐trophic level food webs support the safety of a biocontrol agent 3 years after release

Biological control (biocontrol) is a powerful tool for managing invasive alien species and assisting the restoration of native ecosystems. Rigorous post‐release monitoring of biocontrol agents is critical to evaluate the success of biocontrol programs; however, this is still rarely implemented. Here, we combined the use of species interaction networks with a Before‐After Control‐Impact design to evaluate the target and non‐target, direct and indirect effects of the Australian gall wasp Trichilogaster acaciaelongifoliae , released to control the invasive plant Acacia longifolia in Portugal. We compared the structure of plant‐galling insect‐parasitoid food webs before and 3 years after the release of the biocontrol agent. Exhaustive sampling did not detect any non‐target effects, either direct (on non‐target plants) or indirect (on other galling insects via shared plants). Additionally, no significant changes were detected in network structure that could be related to the establishment of the biocontrol agent. This study shows that monitoring biocontrol at the community level is possible and that, when carefully planned, biocontrol poses minimal risk of non‐target effects.

López‐Núñez, Francisco A. [Centre for Functional E

ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement

Agentic Retrieval-Augmented Code Synthesis with Iterative RefinementIn supercomputing, efficient and optimized code generation is essential to leverage high-performance systems effectively. We have developed Agentic Retrieval-Augmented Code Synthesis (ARCS), an advanced framework for accurate, robust, and efficient code generation, completion, and translation. ARCS integrates Retrieval-Augmented Generation (RAG) with Chain-of-Thought (CoT) reasoning to systematically break down and iteratively refine complex programming tasks. An agent-based RAG mechanism retrieves relevant code snippets, while real-time execution feedback drives the synthesis of candidate solutions. This process is formalized as a state-action search tree optimization, balancing code correctness with editing efficiency. Evaluations on the Geeks4Geeks and HumanEval benchmarks demonstrate that ARCS significantly outperforms traditional prompting methods in translation and generation quality. By enabling scalable and precise code synthesis, ARCS offers transformative potential for automating and optimizing code development in supercomputing applications, enhancing computational resource utilization

Bhattarai, Manish [Los Alamos National Labs]

CodeScribe Agent

SF-26-086 CodeScribe introduces a structured, multi-stage pipeline that combines deterministic program analysis with LLM-powered translation to enable incremental, testable Fortran-to-C++ migration. First, `code-scribe index` traverses the project directory tree and produces `scribe.yaml` metadata files recording all modules, subroutines, and functions at each level, giving the LLM accurate structural context instead of a hallucinated codebase model. Second, `code-scribe draft` performs the deterministic portion of translation — converting Fortran types to C++ equivalents, replacing `use` statements with `#include` and `using namespace` directives, and detecting constructs requiring special handling — while embedding`scribe-prompt` annotations that guide the LLM through non-trivial cases such as statement-function-to-lambda conversions and `extern "C"` wrapper generation. Third, `code-scribe translate` applies project-specific TOML-based few-shot prompt templates and submits the composed prompt to a pluggable LLM backend (OpenAI, Anthropic, Argonne ARGO, any OpenAI-compatible endpoint, or local Hugging Face checkpoints), producing a C++ source file, a header, and a Fortran-C++ interface file for each translated routine so the codebase compiles and runs correctly throughout the migration. Beyond translation, CodeScribe includes a tool-using coding agent (`code-scribe agent`) with read, bash, edit, and write capabilities, and a bounded loop mode (`code-scribe loop`) that runs repeated stateless agent sessions over a task file with restricted tool access — enabling sustained, auditable software development workflows for broader scientific computing tasks.

Dhruv, Akash [Argonne National Laboratory (ANL), A

National Security Programs - Cyber: MMAREJBLIGE – Modular Multi Agent Grid Emulation for Joined Breakdowns in Linked Generative Emulations - 23-0644

Modular Multi Agent Grid Emulations for Joined Breakdowns in Linked Generative Emulations (MMAREJBLIGE) introduces an agent-based modeling framework into real-time cyber-physical emulation to achieve a context-aware environment that introduces operator/attacker/external-condition variability to improve emulation fidelity and testing rigor. We detail our agent framework design, internal communication via message passing, and time synchronization, as well as the individual components of the system. We include a brief analysis of several scenarios run on a real-time, hardware-in-the-loop, Industrial Control Systems (ICS) test-bed which include normal operation, physical disruption, disruption with mitigation, and disruption with mitigation during a cyber denial-of-service (DOS) attack.

42 ENGINEERING

Electrodeposition of Tungsten using Hydrotropic Agents

The current work sought to electrodeposit tungsten from water-based solutions. The work’s initial hypothesis was that methoxide reducing agents could be used to generate urea anions that could enable tungsten electrodeposition. Unfortunately, this hypothesis was found to be incorrect. However, the work led to the understanding that chemical reducing agents not only enable the electrodeposition of refractory metals like rhenium from water-based solutions, but that the key to plating tungsten in the future involves chemical reduction followed by stabilization with proper ligands. The work resulted in a manuscript under review at Inorganic Chemistry Communications on the discovered of L-histidine as a suitable reducing agent from rhenium electrodeposition. Rhenium-tungsten alloys with 4% tungsten were deposited. A technical advance was filed for the rhenium chemistry and parts were delivered to an internal Sandia customer that used the chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Expanding Activity Allocation Models to Daily Activities: Tracking Simulated Agent Trips to Exercise Locations in Clarksville, TN

Exercise facilities have been proven to have numerous physical, mental, and psychological benefits, yet exercise facilities are still inaccessible to a large portion of the population. This study serves to explore the accessibility of fitness centres through geographical, demographic, and temporal lenses through an expansion of the UrbanPop framework that seeks to allocate simulated agents to fitness centres in the Clarksville metro to explore the effects of travel distances on different demographics throughout the week. Findings indicate that senior and retired demographics consistently travel longer distances to exercise in the larger Clarksville area, likely due to tendencies to live further from the center of the metropolitan area. Furthermore, all demographics tend to travel further distances to exercise on the weekends rather than the weekdays, indicating that travel distance can affect likelihood of agents to travel, especially on weekdays when many agents are in the workforce or participating in schooling.

99 GENERAL AND MISCELLANEOUS

Exploring Large Language Model Agents in Cybersecurity: A Literature Review with Experiments

The accelerated development and integration of large language model (LLM) agents have led researchers and developers to explore their effectiveness in cybersecurity, specifically with penetration testing (pentesting). Recent research efforts have attempted to use LLM agents to automate the process of pentesting because of the cost and time requirements that are required to perform a manual review. However, not all of the tools perform as expected. This paper reviews some of the newest and most popular autonomous pentesting frameworks, highlighting the capabilities and limitations of each one with the goal of providing the components needed to successfully and effectively build an autonomous pentesting agent in the future.

97 MATHEMATICS AND COMPUTING

MSD CoP Webinar: "Generative agents: A new frontier for representing human actors and their behavior in MSD models"

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Talk #1: Behavioral Generative Agents for Energy Operations Presenter: Dr. Cong Chen (Thayer School of Engineering, Dartmouth College) Abstract: Accurately modeling consumer behavior in energy operations remains challenging due to inherent uncertainties, behavioral complexities, and limited empirical data. This talk introduces a novel approach leveraging generative agents--artificial agents powered by large language models--to realistically simulate customer decision-making in dynamic energy operations. Talk #2: Simulating multiple human perspectives in socio-ecological systems using large language models Presenter: Dr. Yongchao Zeng (Institute of Meteorology and Climate Research, Atmospheric Environmental Research (IMK-IFU) of the Karlsruhe Institute of Technology in Germany) Abstract: Understanding socio-ecological systems requires insights from diverse stakeholder perspectives. This talk describes a novel simulation system called HoPeS (Human-oriented Perspective Shifting). HoPeS enables model users to not only explore simulated socio-ecological systems (SESs) from a third-person observer's perspective but also take any of the simulated stakeholder roles, like playing an RPG game. By shifting multiple perspectives, model users can reflect and integrate the situated knowledge learned through the participatory simulation, approximating a more holistic and less biased understanding of SESs. Moderators: Jim Yoon (MSD CoP Human Systems Modeling Working Group Co-Chair); Stefano Galelli (MSD CoP Using AI to Enhance MSD Research Working Group Co-Chair); Patrick M. Reed (MSD CoP Facilitation Team) This webinar was held on: November 13th, 2025 from 12-1 PM EST.

Artificial Intelligence

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]

Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents

Membership inference attacks (MIAs), which enable adversaries to determine whether specific data points were part of a model's training dataset, have emerged as an important framework to understand, assess, and quantify the potential information leakage associated with machine learning systems. Designing effective MIAs is a challenging task that usually requires extensive manual exploration of model behaviors to identify potential vulnerabilities. In this paper, we introduce AutoMIA -- a novel framework that leverages large language model (LLM) agents to automate the design and implementation of new MIA signal computations. By utilizing LLM agents, we can systematically explore a vast space of potential attack strategies, enabling the discovery of novel strategies. Our experiments demonstrate AutoMIA can successfully discover new MIAs that are specifically tailored to user-configured target model and dataset, resulting in improvements of up to 0.18 in absolute AUC over existing MIAs. This work provides the first demonstration that LLM agents can serve as an effective and scalable paradigm for designing and implementing MIAs with SOTA performance, opening up new avenues for future exploration.

Tran, Toan Viet [Emory University]

Bayesian Calibration of Stochastic Agent Based Model via Random Forest

Agent-based models (ABM) provide an excellent framework for modeling outbreaks and interventions in epidemiology by explicitly accounting for diverse individual interactions and environments. However, these models are usually stochastic and highly parametrized, requiring precise calibration for predictive performance. When considering realistic numbers of agents and properly accounting for stochasticity, this high-dimensional calibration can be computationally prohibitive. This paper presents a random forest-based surrogate modeling technique to accelerate the evaluation of ABMs and demonstrates its use to calibrate an epidemiological ABM named CityCOVID via Markov chain Monte Carlo (MCMC). The technique is first outlined in the context of CityCOVID's quantities of interest, namely hospitalizations and deaths, by exploring dimensionality reduction via temporal decomposition with principal component analysis (PCA) and via sensitivity analysis. The calibration problem is then presented, and samples are generated to best match COVID-19 hospitalization and death numbers in Chicago from March to June in 2020. Further, these results are compared with previous approximate Bayesian calibration (IMABC) results, and their predictive performance is analyzed, showing improved performance with a reduction in computation.

60 APPLIED LIFE SCIENCES

Towards philosophical reasoning with agentic LLMs: Socratic method for scientific assistance

As large language models (LLMs) become central tools in science, improving their reasoning capabilities is critical for meaningful and trustworthy applications. We introduce a Socratic agent for scientific reasoning, implemented through a structured system prompt that guides LLMs via classical principles of inquiry. Unlike typical prompt engineering or retrieval-based methods, our approach leverages definition, analogy, hypothesis elimination, and other Socratic techniques to generate more coherent, critical, and domain-aware responses. We evaluate the agent across diverse scientific domains and benchmark it on the abstraction and reasoning corpus challenge dataset, achieving 97.15% under a fixed prompting protocol and without fine-tuning or external tools. Expert evaluation shows improved reasoning depth, clarity, and adaptability over conventional LLM outputs, suggesting that structured prompting rooted in philosophical reasoning can improve the scientific utility of language models.

LLM reasoning

Biphasic response of human iPSC-derived neural network activity following exposure to a sarin-surrogate nerve agent

Organophosphorus nerve agents (OPNA) are hazardous environmental exposures to the civilian population and have been historically weaponized as chemical warfare agents (CWA). OPNA exposure can lead to several neurological, sensory, and motor symptoms that can manifest into chronic neurological illnesses later in life. There is still a large need for technological advancement to better understand changes in brain function following OPNA exposure. The human-relevant in vitro multi-electrode array (MEA) system, which combines the MEA technology with human stem cell technology, has the potential to monitor the acute, sub-chronic, and chronic consequences of OPNA exposure on brain activity. However, the application of this system to assess OPNA hazards and risks to human brain function remains to be investigated. In a concentration-response study, we have employed a human-relevant MEA system to monitor and detect changes in the electrical activity of engineered neural networks to increasing concentrations of the sarin surrogate 4-nitrophenyl isopropyl methylphosphonate (NIMP). We report a biphasic response in the spiking (but not bursting) activity of neurons exposed to low (i.e., 0.4 and 4 μM) versus high concentrations (i.e., 40 and 100 μM) of NIMP, which was monitored during the exposure period and up to 6 days post-exposure. Regardless of the NIMP concentration, at a network level, communication or coordination of neuronal activity decreased as early as 60 min and persisted at 24 h of NIMP exposure. Once NIMP was removed, coordinated activity was no different than control (0 μM of NIMP). Interestingly, only in the high concentration of NIMP did coordination of activity at a network level begin to decrease again at 2 days post-exposure and persisted on day 6 post-exposure. Notably, cell viability was not affected during or after NIMP exposure. Also, while the catalytic activity of AChE decreased during NIMP exposure, its activity recovered once NIMP was removed. Gene expression analysis suggests that human iPSC-derived neurons and primary human astrocytes resulted in altered genes related to the cell’s interaction with the extracellular environment, its intracellular calcium signaling pathways, and inflammation, which could have contributed to how neurons communicated at a network level.

59 BASIC BIOLOGICAL SCIENCES