Search NASA⌕ Search

SEARCH · Search NASA

Results for “knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

GeoBridge: Unearthing Insights from Connecting Communities to Geothermal Information and Opportunities: Preprint

Knowledge is essential for overcoming obstacles in the development and adoption of geothermal technologies, and the geothermal community is home to numerous tools, events and organizations dedicated to sharing knowledge. However, many of these tools can be difficult to find, their resources undiscoverable by search engines, available only to members, or hidden away behind pay walls (Weers et al., 2024). The Department of Energy's (DOE) GeoBridge was developed by the National Renewable Energy Laboratory (NREL) to help bridge gaps in information and connect the geothermal community to the resources it needs. Launched in October 2024, GeoBridge aspires to expand the pool of geothermal stakeholders by providing in-roads to geothermal information, tools, and community resources. It helps to make these resources available to the broader geothermal community as well as those looking to join, such as entrepreneurs or innovators in adjacent industries looking to expand into geothermal energy. This paper explores a post-launch analysis of GeoBridge including data from analytics, feedback from GeoBridge users, the geothermal community, and the GeoBridge Advisory Group as well as an analysis of efficacy of various promotions for GeoBridge.

15 GEOTHERMAL ENERGY↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

On the Recoverability of Reactor Dynamics from Point Kinetics Data using SINDYc

Data driven models for reactor dynamics tend to face a few notable challenges. First, the broad range of timescales present across the various feedback mechanisms. Next, high standards for safety and importance of model performance in all operational domains. Finally, the correlations between state variables observed in most transients leads to difficulties when methods attempt to attribute certain dynamic phenomena to a particular cause. The current paper seeks to support efforts towards incorporating physics knowledge into one particular data-driven method for finding reactor dynamics called ``Sparse Identification of Nonlinear Dynamics with Control (SINDYc). The incorporation of physical knowledge into SINDYc will help address the challenges listed above by giving the model an initial understanding of the system being modeled. The demonstration of an exact representation of a set of point reactor dynamics equations in SINDYc is provided. Then, this allows for further discussion with mathematical justification as to why SINDYc, and other data-driven methods, may be difficult to apply to nuclear reactor dynamics due to correlations in the state variables. A particularly useful result of this work is the set of candidate functions required in SINDYc to exactly represent the reactor dynamics under a point approximation. In the future, SINDYc can be integrated with point models for reactor dynamics before being applied to the physical system to yield higher accuracy.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Agentic Diagrammatica: Towards Autonomous Symbolic Computation in High Energy Physics

We present Diagrammatica, a symbolic computation extension to the HEPTAPOD agentic framework, which enables LLM agents to plan and execute multi-step theoretical calculations. Symbolic computation poses a distinctive reliability challenge for LLM agents, as correctness is governed by implicit mathematical conventions that are not encoded in a form that can be easily checked in the computational backend. We identify two complementary remedies, tool-constrained computation and targeted knowledge grounding, and pursue the first as the primary architecture. Concretely, we concentrate the agent's action distribution onto tool calls with convention-fixing semantics, in which the agent specifies a compact, human-auditable diagram specification and a trusted backend performs the symbolic or numerical manipulations exactly. The toolkit provides two complementary calculation paths consuming a shared diagram specification: Naive Dimensional Analysis (NDA) for order-of-magnitude rate estimates and Exact Diagrammatic Analysis (EDA) for tree-level symbolic calculations via automatic FeynCalc code generation, both supplemented by automatic Feynman diagram enumeration and a navigable theory knowledge base. The architecture is validated on two benchmarks: (1) an exhaustive catalog of all tree-level, single-vertex $1\to 2$ partial decay widths across scalar, fermion, and vector parents, with complete massless and threshold limits and Standard Model validation; and (2) an NDA sensitivity study of the muon decay multiplicity $μ^+ \to ν_μ\barν_e + n(e^+e^-) + e^-$, determining the maximum observable $n$ at current and planned muon experiments.

Menzo, Tony [Alabama U.; Fermilab] (ORCID:00000002↗

Understanding and design of interstitial oxygen conductors

Highly efficient oxygen-active materials that react with, absorb, and transport oxygen is essential for fuel cells, electrolyzers and related applications. While vacancy-mediated oxygen-ion conductors have long been the focus of research, they are limited by high migration barriers at intermediate temperatures (400–600 °C), which hinder their practical applications. In contrast, interstitial oxygen conductors exhibit significantly lower migration barriers enabling higher ionic conductivity at lower temperatures. This review systematically examines both well-established and recently identified families of interstitial oxygen-ion conductors, focusing on how their unique structural motifs such as corner-sharing polyhedral frameworks, isolated polyhedral, and cage-like architectures, facilitate low migration barriers through interstitial and/or interstitialcy diffusion mechanisms. A central discussion of this review focuses on the evolution of design strategies, from targeted donor doping, element screening, to physical-intuition descriptor material screening and machine learning approach, which leverage computational tools to explore vast chemical spaces in search for new interstitial conductors. The success of these strategies demonstrates that a significant, largely unexplored space remains for discovering high-performing interstitial oxygen conductors. Crucial features enabling high-performance interstitial oxygen diffusion include the availability of electrons for oxygen reduction and sufficient structural flexibility with accessible volume for interstitial accommodation and migration. This review concludes with a forward-looking perspective, proposing a knowledge-driven methodology that integrates current understanding with data-centric approaches to identify promising interstitial oxygen conductors outside traditional search paradigms. These approaches are expected to significantly accelerate the development of high-performance interstitial oxygen conductors for a variety of oxygen-active applications, ultimately paving the way for more efficient and sustainable energy technologies.

Interstitial oxygen conductors↗

Towards the next generation of Geospatial Artificial Intelligence

Geospatial Artificial Intelligence (GeoAI), as the integration of geospatial studies and AI, has become one of the fastest-developing research directions in spatial data science and geography. This rapid change in the field calls for a deeper understanding of the recent developments and envision where the field is going in the near future. In this work, we provide a quantitative analysis of the GeoAI literature from the spatial, temporal, and semantic aspects. We briefly discuss the history of AI and GeoAI by highlighting some pioneering work. Then we discuss the current landscape of GeoAI by selecting five representative subdomains including remote sensing, urban computing, Earth system science, cartography, and geospatial semantics. Finally, we highlight several unique future research directions of GeoAI which are classified into two groups: GeoAI method development challenges and GeoAI Ethics challenges. Topics include heterogeneity-aware GeoAI, knowledge-guided GeoAI, spatial representation learning, geo-foundation models, fairness-aware GeoAI, privacy-aware GeoAI, as well as interpretable and explainable GeoAI. We hope our review of GeoAI’s past, present, and future is comprehensive and can enlighten the next generation of GeoAI research.

58 GEOSCIENCES↗

Language models for materials discovery and sustainability: Progress, challenges, and opportunities

Significant advancements have been made in one of the most critical branches of artificial intelligence: natural language processing (NLP). These advancements are exemplified by the remarkable success of OpenAI’s GPT-3.5/4 and the recent release of GPT-4.5, which have sparked a global surge of interest akin to an NLP gold rush. Here, in this article, we offer our perspective on the development and application of NLP and large language models (LLMs) in materials science. We begin by presenting an overview of recent advancements in NLP within the broader scientific landscape, with a particular focus on their relevance to materials science. Next, we examine how NLP can facilitate the understanding and design of novel materials and its potential integration with other methodologies. To highlight key challenges and opportunities, we delve into three specific topics: (i) the limitations of LLMs and their implications for materials science applications, (ii) the creation of a fully automated materials discovery pipeline, and (iii) the potential of GPT-like tools to synthesize existing knowledge and aid in the design of sustainable materials.

36 MATERIALS SCIENCE↗

A Privacy-Preserving Cyber Threat Intelligence Sharing System

Cyber Threat Intelligence (CTI) is a key resource for developing defensive strategies against potential cyber adversaries. Entities typically access CTI through open-source platforms, national agencies, or specialized commercial services. However, the bi-directional exchange of CTI is hindered by organizational trust boundaries, which complicate the sharing processes between entities and CTI providers. Centralized CTI services benefit from receiving suspicious cyber observables such as IP addresses, domain names, and email addresses from various entities. The aggregation allows for the correlation of widespread adversarial activities to enhance the alert and response mechanisms across the network of involved parties. Despite these benefits, openly sharing such observables incurs potential legal, regulatory, and reputational risks for the disclosing entities.This paper introduces a system designed to facilitate the secure exchange of cyber observables across trust boundaries without compromising the anonymity of the sharing entities. Here, we propose an architecture that leverages common web protocols alongside zero-knowledge proofs to authenticate members while maintaining anonymity. Additionally, we outline a privacy model tailored for STIX (Structured Threat Information eXpression) cyber observables to minimize the risk of inadvertently disclosing private information. Through our threat models, we assess the privacy implications of our proposed system and demonstrate its potential to enhance collaborative cyber defense efforts without exposing entities to undue risk.

BBS+ Signatures↗

Cryogen Safety Live #8876

Cryogenics (from the Greek word κρυος, meaning frost or icy cold) is the study of the effects and behavior of materials at very low temperature. This course is designed to provide trainees with an introduction to cryogen use, hazards associated with cryogen systems, cryogen safety components, and the requirements that govern the design and use of cryogen systems at Los Alamos National Laboratory (LANL). The knowledge you gain is intended to help you keep your workplace safe for you and your coworkers.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Connecting Minds: AI Use Cases to Bridge Power Systems and Large Language Models for Practical Applications

Recent advances in artificial intelligence (AI) and development of large language models (LLMs) present the opportunity to develop a new generation of power systems applications. In contrast with early power system AI applications based on structured numerical data, LLMs offer unique capabilities to perform logical reasoning using text documents, unstructured data, and application programming interface (API) calls to computational software. This paper seeks to bridge the knowledge gap between power systems engineers and LLM developers through a crosscutting explanation of use cases, characteristics, requirements, practical considerations from the perspectives of both LLM capabilities and industry needs. Specific focus is given to applications that can be realistically deployed by electric utilities. After introducing the architecture of LLMs and unique challenges of the power systems domain, this paper proposes twenty representative LLM applications grouped into categories of 1) power system operations, 2) asset management, 3) system planning and analytics, and 4) energy management and protection systems. Five use cases are presented within each category with descriptions of the motivation, objectives, approaches, example inputs / outputs, and benefits of each use case.

24 POWER TRANSMISSION AND DISTRIBUTION↗

National Hydropower Fish Passage Database User's Guide and Methodology

Fish passage facilities are used to mitigate the impacts of hydropower dams on migratory fish in rivers, but information on the location, types, and characteristics of this infrastructure is incomplete at a national scale. Researchers at Oak Ridge National Laboratory partnered with federal agency and industry stakeholders to create the first national-scale database of fish passage infrastructure at U.S. hydropower developments. This database contains information of great value to a broad range of stakeholders; it includes fish passage facility engineering characteristics, targeted fish species, operational schedules, and costs. This data resource addresses a large gap in knowledge of the deployment of fish passage technology and is freely available to members of the hydropower community, including federal and state regulators and resource agencies, non-governmental organizations, industry, and other user groups to support project planning and regulatory (re)licensing activities. This database supports the U.S. Department of Energy Water Power Technologies Office objective to develop decision support tools and data resources that improve environmental performance and ensure hydropower’s long-term value to the American public.

13 HYDRO ENERGY↗

What Is the Agent Doing? Visualizing Agentic AI Querying Workflows

We explore how visualizations can help users understand what an AI agent is doing as it builds and runs queries over data. As part of the LinkQ system, a natural language interface for querying knowledge graphs with a large language model (LLM), we designed two complementary views: A State Diagram that shows where the agent is within a larger workflow, and a Live Action Display that gives real-time updates about the agent's current task. In a study with 14 practitioners, we found that these visuals helped participants build stronger mental models of the agent's behavior while also increasing their confidence in the system. However, we also observed that users sometimes trusted incorrect outputs simply because the agent appeared to be doing the "right" thing. Our findings point to both the value and risk of visualizing agent behavior in interactive AI systems.

97 MATHEMATICS AND COMPUTING↗

MCNP® Code Version 6.3.2 Theory & User Manual (Rev. 1)

This document acts as a repository of knowledge for the Monte Carlo N-Particle (MCNP) transport computer code. It is maintained alongside the source code and attempts to introduce new users and re-familiarize experienced users with the theory and practices of using the MCNP code for the wide range of particle transport analyses that it is appropriate for. The latest version of the MCNP code, version 6.3.2, provides the Monte Carlo particle transport community with the latest feature developments and bug fixes in the MCNP code. The MCNP code version 6.0 and later is also known as the MCNP6 code.

42 ENGINEERING↗

Computational tools and data integration to accelerate vaccine development: challenges, opportunities, and future directions

The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.

60 APPLIED LIFE SCIENCES↗

Mapping Hsp104 interactions using cross‐linking mass spectrometry

Molecular machines from the AAA+ (ATPases Associated with diverse cellular Activity) superfamily of protein disaggregases play important roles in protein folding, disaggregation and DNA processing. Recent cryo-EM structures of AAA+ molecular machines have uncovered nuanced changes in their conformation that underlie their specialized functions. Structural knowledge of these molecular machines in complex with substrates begins to explain their mechanism of activity. Here, we explore how cross-linking mass spectrometry (XL-MS) can be used to interpret changes in conformation induced by ATP in Hsp104 and how a substrate may interact with Hsp104. We applied a panel of cross-linking reagents to produce cross-linking maps of Hsp104 and interpret our data on previously determined X-ray and cryo-EM structures of Hsp104 from a thermophilic yeast, Calcarisporiella thermophila. We developed an analysis pipeline to differentiate between intra-subunit and inter-subunit contacts within the hexameric homo-oligomer. We identify cross-links that break the asymmetry that is present in Hsp104 in an ATP-hydrolysis competent conformation but is absent in an ATP-hydrolysis-defective mutant. Finally, we identify contacts between Hsp104 and a selected protein (proprotein convertase subtilisin/kexin type 9 PCSK9) to reveal contacts on the central channel of Hsp104 across the length of this protein indicating that we might have trapped interactions consistent with its translocation. Our simple and robust XL-MS-based experiments and methods help interpret how these molecular machines change conformation and bind to other proteins even in the context of homo-oligomeric assemblies enabling coupling state-of-the-art modeling approaches with XL-MS.

60 APPLIED LIFE SCIENCES↗

Physics-informed Deep Reinforcement Learning-based Control in Power systems

Incorporating physics information into the deep reinforcement learning (DRL) process is a promising approach for addressing the challenges faced in learning-based control design problems for physical systems. Power grid dynamics, being a physical system, adheres to specific physical laws, constraints, as well as operational and control rules. Therefore, consideration of such physics-based law improves the learning process drastically. In general, traditional grid control schemes rely on rule-based mechanisms that cannot adapt to changing operating conditions. To improve the adaptability and computation time, recent research has seen a surge of DRL-based applications in power grid control. A generic DRL-based control design imposes the system performance requirements through the design of reward functions. In some cases, some of the important physics information is injected through this reward function. However, due to the complex dynamics and large state-action space, learning an optimal DRL policy often becomes challenging. Inspired by the latest developments in general machine learning (ML) research, power system researchers have been investigating more direct ways of incorporating physics knowledge into DRL training. This chapter specifically focuses on these aspects of physics-informed DRL designs in grid control. It discusses the significance, applications, research gaps, and open problems that need to be addressed in future research.

artificial intelligence, machine learning↗

Unveiling Atomistic Mechanisms Governing Additive Manufacturing Processability and Mechanical Behavior of a Refractory Complex Concentrated Alloy

Extending the concept of complex concentrated alloys (CCAs) to the refractory alloys (solidus temperature over 2000 °C) space potentially facilitates the design of lightweight structural alloys with service temperatures that exceed those of Ni and Co‐based alloys. However, the room and elevated temperature tensile properties of the current refractory‐CCAs (R‐CCAs) are inferior to those of the Ni/Co‐based alloys. Furthermore, the manufacturing scalability of R‐CCAs remains challenging, in that cracks are prevalent in all R‐CCAs when processed using near‐net shape manufacturing processes, such as fusion‐based additive manufacturing (F‐BAM). Still, mechanisms governing the poor F‐BAM processability of R‐CCAs remain unexplored. Here, to this end, this work unveils the atomistic mechanisms underlying F‐BAM process‐induced cracking in a NbTiTaMoHfZrC R‐CCA. The implications of light elements’ presence for intrinsic ductility and grain boundary cohesion, and subsequently for F‐BAM processability and mechanical behavior, are revealed. Leveraging the insights, we accomplish what is, to the best of the knowledge, the first instance of crack‐free F‐BAM processing of any R‐CCA. Additionally, the R‐CCA exhibits over 20% tensile ductility and ≈160 MPa tensile yield strength at 1200 °C. In addition to facilitating the design of lightweight R‐CCAs, findings enable scalable manufacturing of these ultra‐high temperature alloys for structural applications.

Refractory alloys↗