Search NASASearch

SEARCH · Search NASA

Results for “Knowledge Modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Impacts of Biomass Feedstock Pre-Processing on Heat and Mass Transfer During Pyrolysis Using X-Ray Computed Tomography and Multiscale Modeling

Knowledge of the transport properties of biomass particles such as porosity, tortuosity, and permeability is paramount for high-fidelity modeling of biomass pyrolysis due to the heat and mass transfer limitations imposed by particle microstructure. X-ray computed tomography (XCT) is a non-destructive imaging method that enables full 3D reconstructions of the biomass particle microstructure with high resolution, permitting direct calculation of porosity, tortuosity, and permeability from real particle geometries. In this study, XCT imaging revealed the 3D microstructures of particles and chars from pyrolytic conversion of cylindrically cut or milled/pelletized loblolly pine samples. The porosity, tortuosity, and permeability were calculated directly from the XCT geometries via open-source microstructural analysis tool MATBOX+TauFactor (https://github.com/NREL/MATBOX_Microstructure_analysis_toolbox) and computational fluid dynamics (CFD) simulations using our solver, Mesoflow (https://github.com/NREL/mesoflow). These properties were used in a reactor scale model developed in COMSOL of the single particle reactor at NREL to investigate the impact of feedstock pre-processing on biomass conversion during pyrolysis with rigorous experimental validation.

biomass

CHEMREASONER: Heuristic Search over a Large Language Model’s Knowledge Space using Quantum-Chemical Feedback

The discovery of new catalysts is essential for the design of new and more efficient chemical processes in order to transition to a sustainable future. We introduce an AI-guided computational screening framework unifying linguistic reasoning with quantum-chemistry based feedback from 3D atomistic representations. Our approach formulates catalyst discovery as an uncertain environment where an agent actively searches for highly effective catalysts via the iterative combination of large language model (LLM)-derived hypotheses and atomistic graph neural network (GNN)-derived feedback. Identified catalysts in intermediate search steps undergo structural evaluation based on spatial orientation, reaction pathways, and stability. Scoring functions based on adsorption energies and barriers steer the exploration in the LLM's knowledge space toward energetically favorable, high-efficiency catalysts. We introduce planning methods that automatically guide the exploration without human input, providing competitive performance against expert-enumerated chemical descriptor-based implementations. By integrating language-guided reasoning with computational chemistry feedback, our work pioneers AI-accelerated, trustworthy catalyst discovery.

artificial intelligence

Large Language Model Integration for Knowledge Retrieval and Interaction for the DUNE Experiment

The Deep Underground Neutrino Experiment (DUNE) is a next-generation neutrino experiment that will generate an unprecedented volume of heterogeneous information-from documentation and technical notes to experimental data and reconstruction pipelines. Efficient knowledge retrieval and contextual understanding are increasingly critical for collaboration-wide productivity and onboarding. In this work, we present DUNE-GPT, a prototype framework that leverages large language models (LLMs) and retrieval-augmented generation (RAG) to enable natural-language querying of DUNE's internal documentation and technical resources. The system provides an intelligent interface for DUNE collaborators to interact with experiment-specific knowledge while maintaining data privacy and infrastructure compliance within Fermilab computing resources.

Rafique, A. [Argonne (main)]

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference

A knowledge-informed large language model framework for U.S. nuclear power plant shutdown initiating event classification for probabilistic risk assessment

Identifying and classifying shutdown initiating events (SDIEs) is critical for developing shutdown probabilistic risk assessment for nuclear power plants. Existing computational approaches cannot achieve satisfactory performance due to the challenges of unavailable large, labeled datasets, imbalanced event types, and label noise. To address these challenges, we propose a hybrid pipeline that integrates a knowledge-informed machine learning model to prescreen non-SDIEs and a large language model (LLM) to classify SDIEs into four types. In the prescreening stage, we proposed a set of 44 SDIE text patterns that consist of the most salient keywords and phrases from six SDIE types. Text vectorization based on the SDIE patterns generates feature vectors that are highly separable by using a simple binary classifier. The second stage builds Bidirectional Encoder Representations from Transformers (BERT)-based LLM, which learns generic English language representations from self-supervised pretraining on a large dataset and adapts to SDIE classification by fine-tuning it on an SDIE dataset. The proposed approaches are evaluated on a dataset with 10,928 events using precision, recall ratio, F 1 score, and average accuracy. In conclusion, the results demonstrate that the prescreening stage can exclude more than 97% non-SDIEs, and the LLM achieves an average accuracy of 95.1% for SDIE classification.

99 - GENERAL AND MISCELLANEOUS

Classification of compact objects and model comparison using EOS knowledge

Nuclear theory and experiments, alongside astrophysical observations, constrain the equation of state (EOS) of supranuclear-dense matter. Conversely, knowledge of the EOS allows an improved interpretation of nuclear or astrophysical data. In this article, we use several established constraints on the EOS and the new NICER measurement of PSR J0437-4715 to comment on the nature of the primary companion in GW230529 and the companion of PSR J0514-4002E. We find that, with a probability of ≳84% and ≳68%, respectively, both objects are black holes. These likelihoods increase to above 95% when one uses GW170817’s remnant as an upper limit on the TOV mass. We also demonstrate that the current knowledge of the EOS substantially disfavors high masses and radii for PSR J⁢0030+0451, inferred recently when combining NICER with XMM-Newton background data and using particular hot-spot models. Lastly, we also use our obtained EOS knowledge to comment on measurements of the nuclear symmetry energy, finding that the large value predicted by the PREX-II measurement displays some mild tension with other constraints on the EOS.

79 ASTRONOMY AND ASTROPHYSICS

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)

Knowledge Graph Entity Linking using Graph Embeddings

Details the use of a custom embedding model on knowledge graphs to aid in downstream natural language processing (NLP) models for Derivative Classification Assist. Motivations, algorithms, and results were discussed.

Mahesh, Aarav [Sandia National Laboratories (SNL-N

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ

Model-Based Approaches to Generate Knowledge from Data in a Plant Reliability Context

One challenge that nuclear power plant system engineers are facing is continuous generation of an extremely large amount of equipment reliability (ER) data. These data elements come in textual (e.g., condition reports) and numeric (e.g., generated by monitoring systems) forms. They provide system engineers with valuable insights and information by discovering anomalous behaviors or degradation trends, identifying possible causes behind such behaviors and trends, and predicting their direct consequences. This paper directly targets the knowledge generation from ER data by putting “data into context.” We employ model-based system engineering (MBSE) of systems and assets to represent and capture their architecture and functional (i.e., cause-effect) relations. ER data elements are processed by first identifying which of the developed MBSE elements they are referring to. This task is harder for textual data since the information contained in issue or maintenance reports needs to be “understood” by a computational tool. We called this process “knowledge extraction” since our methods extract knowledge from textual data. Last, once numeric and textual ER data elements have been processed and “understood,” we discover possible cause-effect relations among them. This is performed by observing whether a logical connection through the MBSE models exists, and if there is a temporal relationship among them. The logic and temporal are the two main ingredients to perform “machine reasoning” from ER data.

97 - MATHEMATICS AND COMPUTING

Leveraging transfer learning and leaf spectroscopy for leaf trait prediction with broad spatial, species, and temporal applicability

Accurate and reliable prediction of leaf traits is crucial for understanding plant adaptations to environmental variation, monitoring terrestrial ecosystems, and enhancing comprehension of functional diversity and ecosystem functioning. Currently, various approaches (e.g., statistical, physical models) have been developed to estimate leaf traits through hyperspectral remote sensing and leaf spectroscopy. However, the absence of high-performing, transferable, and stable models across various domains of space, plant functional types (PFTs) and seasons hinder our ability to quantify and comprehend spatiotemporal variations in leaf traits. This study proposes robust and highly transferable models for better predicting leaf traits with hyperspectral reflectance. Initially, three datasets were assembled, pairing common leaf traits — chlorophyll (Chla+b), carotenoids (Ccar), leaf mass per area (LAM), equivalent water thickness (EWT) — with leaf spectra measurements collected across diverse geographic locations in the U.S. and Europe, PFTs, and seasons. Measurements were acquired using spectroradiometers (e.g., ASD FieldSpec 3/4/Pro and SVC HR-1024i) with integrating spheres, leaf clips, and contact probes. Here, we then developed transfer learning-based hybrid models that incorporated the domain knowledge of radiative transfer models (RTMs) through pretraining processes and were well-constrained by fine-tuning with field measurements. Through comparison with other state-of-the-art statistical models, including partial-least squares regression (PLSR) and Gaussian Process Regression (GPR), as well as pure physical models, we found that the proposed transfer learning models achieved better predictive performance and higher transferability. Specifically, compared to other statistical models and pure RTMs, the transfer learning model exhibited higher coefficient of determination (R 2 ) values with range of 0.01 to 0.79, lower normalized root mean square error (NRMSE) with range of 0.06 % to 33.25 % in model performance. Additionally, the models exhibited improved transferability, with higher R 2 values range from 0.04 to 0.32, lower NRMSE range from 0.08 % to 30.81 %. The findings underscore that transfer learning models through integrating domain knowledge from RTMs and limited observations, can harness the advantages of both RTMs and statistical models and serve as a promising approach for effectively predicting leaf traits.

59 BASIC BIOLOGICAL SCIENCES

A Roadmap for a Lightning Modeling Grand Challenge

This document is a roadmap for building an interconnected model of the physical processes that produce a lightning discharge, and its observable optical and radio signals. We call this a Lightning Modeling Grand Challenge, recognizing that significant effort and coordination of human and financial resources is required to realize the capability. The roadmap serves to outline the coordination of resources necessary to enable stitching together existing knowledge and model components to make a lightning prediction, and to test these predictions with observations. Such a capability does not currently exist. The roadmap is motivated not only by a spirit of scientific inquiry, but by practical challenges faced by US Federal and societal stakeholders. Advancements in lightning observations have outpaced our tests of integrated understanding, leaving many stakeholders unsure how to design their missions to properly detect and discriminate lightning, and unsure how to apply the sometimes-disagreeing lightning signals from diverse instruments. The time is right to connect existing theories and models to support stakeholders in understanding the signals they observe, for needs as diverse as climate monitoring, national security, weather forecasting, public safety, and protection of natural and built environments. The roadmap’s two main technical sections describe the components of a linked physical model, followed by a description of models of lightning signals and sensors that are driven by outputs from the physical model. The goal is to predict the time-varying physical properties of lightning that are self-consistent with the thunderstorm’s structure and dynamics. These lightning signals then propagate through the storm, with realistic dispersion and attenuation, to receivers on the ground or in space. At a high level, the model begins with weather (cloud) model output, including explicit prediction of the electrification of cloud particles. The cloud’s electrical structure drives a model of lightning physics, from initiation, through channel development, and discharges along those channels. Key lightning parameters, such as the temperature and currents in the channel, and their space and time distribution, are then used to produce optical and electromagnetic signal sources that propagate to modeled receivers. This architecture therefore generates a dataset suitable for comparison to existing and envisioned observing systems. The need for additional measurements and field campaigns to support model development is described. In each model sub-component, inputs, outputs, uncertainties, evaluation methods, and next steps are summarized, interleaved with references to the scientific literature. Identifying boundaries between the model sub-components aids in segmenting an integrated, complex model into practical work packages and system sub-components, allowing a diverse team to contribute and maintain the system. We estimate that at least five years of effort and a $\$$10M initial investment is necessary to make a significant step forward. Mechanisms to facilitate community coordination, including annual workshops and open-source code repositories, are described.

54 ENVIRONMENTAL SCIENCES

Ontologies for Intelligent Data Science

As anyone even vaguely aware of current technology can tell you, machine learning (ML) and artificial intelligence (AI) have made exceptional breakthroughs in recent years. Generative artificial intelligence (GAI) emerged circa 2022 dominated by Large Language Models (LLMs) and generative tools for images emerged at about the same time.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Insights into Slip-Rate Time Functions, Rupture Parameter Correlations, and Ground Motions from Validated Multicycle Earthquake Ruptures

Earthquake strong-motion predictions using kinematic source modeling require knowledge of the slip-rate functions (SRFs) along the rupture and their distinct characteristics in asperities, background (off-asperity) areas and near the surface. Here, in this study we analyzed SRFs from well-validated, self-consistent, and fully dynamic rupture models from earthquake cycles obeying a rate-and-state friction law, from our companion study (Galvez et al., 2021). The shapes of SRFs in asperities are well described by the regularized Yoffe function (RYF), which has only two parameters: rise time T r and smoothing time T s , which control the generation of long- and short-period ground motions, respectively. In background areas, we demonstrate that, in addition to the primary rupture, multiple secondary ruptures may also nucleate from rupture heterogeneities related to asperities, resulting in SRFs with multiple peaks. Because it is impossible to fit a multiple-peak SRF by the single-peak RYF, we describe SRFs in background areas in an effective way by fitting their amplitude spectra with the RYF spectra. Such spectrally effective RYFs capture salient aspects of seismic-wave generation and can be used in rupture generators for strong motion prediction. We found that small T s values correlate with small characteristic weakening distances, large peak slip rates (PSRs), and large rupture velocities. T r values are larger in background areas and smaller in asperities. Within the shallow aseismic zone, T s values approximately quadruple whereas T r values approximately double. Because of this dominant T s increase, PSR values decrease in the near-surface zone. These features indicate that the generation of strong motions by the near-surface portions of the rupture is negligible in the studied scenarios.

Geosciences

Analysis and optimization of seismic monitoring networks with Bayesian optimal experimental design

SUMMARY Monitoring networks increasingly aim to assimilate data from a large number of diverse sensors covering many sensing modalities. Bayesian optimal experimental design (OED) seeks to identify data, sensor configurations or experiments which can optimally reduce uncertainty and hence increase the performance of a monitoring network. Information theory guides OED by formulating the choice of experiment or sensor placement as an optimization problem that maximizes the expected information gain (EIG) about quantities of interest given prior knowledge and models of expected observation data. Therefore, within the context of seismo-acoustic monitoring, we can use Bayesian OED to configure sensor networks by choosing sensor locations, types and fidelity in order to improve our ability to identify and locate seismic sources. In this work, we develop the framework necessary to use Bayesian OED to optimize a sensor network’s ability to locate seismic events from arrival time data of detected seismic phases at the regional-scale. This framework requires five elements: (i) A likelihood function that describes the distribution of detection and traveltime data from the sensor network, (ii) A prior distribution that describes a priori belief about seismic events, (iii) A Bayesian solver that uses a prior and likelihood to identify the posterior distribution of seismic events given the data, (iv) An algorithm to compute EIG about seismic events over a data set of hypothetical prior events, (v) An optimizer that finds a sensor network which maximizes EIG. Once we have developed this framework, we explore many relevant questions to monitoring such as: how to trade off sensor fidelity and earth model uncertainty; how sensor types, number and locations influence uncertainty; and how prior models and constraints influence sensor placement.

58 GEOSCIENCES

Long-term hydro-economic analysis tool for evaluating global groundwater cost and supply: Superwell v1.1

Abstract. Groundwater plays a key role in meeting water demands, supplying over 40 % of irrigation water globally, with this role likely to grow as water demands and surface water variability increase. A better understanding of the future role of groundwater in meeting sectoral demands requires an integrated hydro-economic evaluation of its cost and availability. Yet substantial gaps remain in our knowledge and modeling capabilities related to groundwater availability, recharge, feasible locations for extraction, extractable volumes, and associated extraction costs, which are essential for large-scale analyses of integrated human–water system scenarios, particularly at the global scale. To address these needs, we developed Superwell, a physics-based groundwater extraction and cost accounting model that operates at sub-annual temporal and at the coarsest 0.5° (≈50 km × 50 km) gridded spatial resolution with global coverage. The model produces location-specific groundwater supply–cost curves that provide the levelized cost to access different quantities of available groundwater. The inputs to Superwell include recent high-resolution hydrogeologic datasets of permeability, porosity, aquifer thickness, depth to water table, recharge, and hydrogeological complexity zones. It also accounts for well capital and maintenance costs, as well as the energy costs required to lift water to the surface. The model employs a Theis-based scheme coupled with an amortization-based cost accounting formulation to simulate groundwater extraction and quantify the cost of groundwater pumping. The result is a spatiotemporally flexible, physically realistic, economics-based model that produces groundwater supply–cost curves. We show examples of these supply–cost curves and the insights that can be derived from them across a set of scenarios designed to explore model outcomes. The supply–cost curves produced by the model show that most (90 %) nonrenewable groundwater in storage globally is extractable at costs lower than USD 0.57 m−3, while half of the volume remains extractable at under USD 0.108 m−3. The global unit cost is estimated to range from a minimum of USD 0.004 m−3 to a maximum of USD 3.971 m−3. We also demonstrate and discuss examples of how these cost curves could be used by linking Superwell's outputs with other models to explore coupled human–environmental system challenges, such as water resources planning and management, or broader analyses of multisectoral feedbacks.

Global Change Analysis Model (GCAM)

Enhancing EV Motor Design Through Knowledge-Based AI and Hierarchical Fuzzy Logic Model

This work presents a novel approach to optimizing electric vehicle motor design through the integration of Knowledge-Based Artificial Intelligence (KB-AI) and Hierarchical Fuzzy Logic. Traditional motor design processes are time-intensive, relying heavily on iterative simulations and domain-specific expertise. These processes are further complicated by the nonlinear relationships between key design parameters. The proposed framework addresses these challenges by systematically encoding expert knowledge from scientific literature into a fuzzy logic system, allowing for the efficient handling of complex design variables. The hierarchical fuzzy logic model reduces computational complexity by decomposing the nonlinear relationships into manageable rule sets while maintaining design accuracy. The proposed methodology was applied to the design of a 100 kW motor, yielding optimal values for key parameters. This resulted in a compact motor design with a volume of 2.2 liters, showcasing the framework’s ability to deliver high-performance, application-specific motor configurations.

Kumar, Praveen [ORNL] (ORCID:0000000291877857)