Search NASA⌕ Search

SEARCH · Search NASA

Results for “relationship extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES↗

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING↗

Automated Generation of Graph-based Cyber Threat Intel

With the advancement of AI technology and tools, specifically in the cybersecurity domain, both cyber defenders and threat actors are continuously adapting the use of these capabilities to expedite their operations. With this phenomenon, threat intelligence that is up to date, refreshable, and has relevant context to a specific threat becomes more and more important as it enables cybersecurity professionals to gain insight into relevant data and relationships to guide their operations. This project enables users to frequently aggregate threat intelligence from various sources, such as vendor vulnerability advisories affecting critical infrastructure, malware reports, and adversary writeups into a centralized, standardized database. The project utilizes the Structured Threat Intelligence eXpression (STIX) for a standardized, shareable threat intelligence data format and Neo4j as a graph database solution to store STIX nodes and relationships. Initial results of the project include datasets of over 8,000 nodes and 20,000 relationships extracted from over 500 data sources that have been released within the past month.

Threat Intelligence↗

MechBERT: Language Models for Extracting Chemical and Property Relationships about Mechanical Stress and Strain

Language models are transforming materials-aware naturallanguage processing by enabling the extraction of dynamic, context-rich information from unstructured text, thus, moving beyond the limitations of traditional information-extraction methods. Moreover, small language models are on the rise because some of them can perform better than large language models (LLMs) when given domain-specific questionanswer tasks, especially about an application area that relies on a highly specialized vernacular, such as materials science. We therefore present a new class of MechBERT language models for understanding mechanical stress and strain in materials. These employ Bidirectional Encoder Representations for transformer (BERT) architectures. We showcase four MechBERT models, all of which were pretrained on a corpus of documents that are textually rich in chemicals and their stress–strain properties and were fine-tuned on question-answering tasks. We evaluated the level of performance of our models on domain-specific as well as general English-language question-answer tasks and also explored the influence of the size and type of BERT architectures on model performance. We find that our MechBERT models outperform BERT-based models of the same size and maintain relevancy better than much larger BERT-based models when tasked with domain-specific question-answering tasks within the stress–strain engineering sector. These small language models also enable much faster processing and require a much smaller fraction of data to pretrain them, affording them greater operational efficiency and energy sustainability than LLMs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

metnet-direct-auxiliary

Code to reveal the relationship between fungal metabolomic outputs and the exogeneous treatments triggering their production. Two routes to extract the relationships: (1) direct route - for known and putative metabolites induced by treatments, and (2) for discovering unknown analytes induced by treatments. The code outputs the corresponding networks and also network-based metrics to rank the metabolomic outputs and treatments.

Meena, MuraliG↗

BNF C-band Scanning ARM Precipitation Radar 2nd Generation (CSAPR-2) Extracted Radar Columns and In-Situ Sensors (RadCLss)

Corrected Moments in Antenna Coordinates (CMAC) calculates quantitative precipitation estimates (QPE) from empirical relationships based on equivalent radar reflectivity factor, specific differ- ential phase, and specific attenuation. To evaluate these empirical relationships, the Extracted Radar Columns and In-Situ Sensors (RADclss) product was developed. Utilizing Py-ART, RAD- clss extracts CMAC radar columns above various ARM and partner locations. These columns are then spatiotemporally synced with in-situ observations at the surface utilizing the Atmospheric data Community Toolkit (ACT; Theisen et al. 2025), allowing direct comparison of radar parameters with rain gauges and laser disdrometers for further investigation.

bankhead↗

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT↗

From tides to seasons: How cyclic tidal drivers and plant physiology interact to affect carbon cycling at the terrestrial-estuarine boundary (Final technical report)

Coastal ecosystems are among the most biologically and biogeochemically active and diverse systems on Earth. Because they act as important linkages between terrestrial ecosystems and the open ocean, their incorporation in Earth system models (ESMs) is critical to predict coastal and global responses to environmental changes. However, they vary greatly in the magnitude of tides and the volume and timing of freshwater input from land, making it challenging to model the major biogeochemical reactions that control productivity and greenhouse gas emissions across coastal terrestrial aquatic interfaces (TAIs). Our overall objective was to improve mechanistic process understanding and modeling of tidal wetland hydro-biogeochemistry in coastal TAIs. We established a new flux tower site (Ameriflux US-PLo) in the oligohaline part of the Parker River to continuously monitor ecosystem-scale carbon fluxes under temporally varying salinity conditions. The site is co-located with long-term monitoring plots of the Plum Island Ecosystems LTER project. We installed wells and redox sensors in the marsh interior and creek bank, established biomass monitoring plots and deployed novel optode sensors in both locations. We used this data to parameterize plant-mediated transport in PFLOTRAN and tested the impact of soil heterogeneity on porewater constituents and gas fluxes. We collected observations of root oxygen release with a novel planar optode system in the field. Flux data collected during the measurement period encompasses a large variation in salinity ranging from drought to record precipitation years. We developed a method to extract functional relationships from the flux data using artificial neural networks, identifying salinity thresholds for CH 4 fluxes. Finally, we are using the coupled ELM-PFLOTRAN model to test the impact of antecedent hydrological conditions on the salinity-CH 4 flux relationship. This grant contributed to the professional development of one postdoc, three research assistants and one graduate student. The sensor data has been shared with external collaborators.

54 ENVIRONMENTAL SCIENCES↗

Investigation of the Performance and Explainability Tradeoffs for Machine-Learning Models for Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants

Predictive maintenance (PdM) has shown great potential for achieving substantial cost savings and enhancing the economic competitiveness of nuclear power plants (NPPs) in today's energy market. Among the different modeling approaches that exist, machine learning (ML) tools in particular have a demonstrated ability to handle high dimensional and multivariate data and to extract hidden relationships within data in industrial environments. While ML methods show great potential, their lack of explainability---especially for black-box models---is a major hurdle to their adoption. Moreover, considering the supposed trade-off between explainability and performance challenges, careful consideration must be made as to which of these quality aspects takes precedence in light of multiple modeling options, resource availability, and domain characteristics. The present work evaluates the performance of six ML models, each with a different degree of explainability, in classifying the conditions of circulating water pumps (CWPs) by utilizing sensor data from nuclear power plants. To determine the drivers behind the trade-offs presented by this array of models, this work also tests different combinations of CWP units as the training and testing data, degrees of data imbalance, and objective functions for hyperparameter tuning. It was found that black-box models tend to afford superior performance in cases where there are far more instances of one type of labeled data than of any other type. It is recommended that a guided procedure be followed for designing and delivering an ML system that is sufficiently explainable to all involved stakeholders.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Measuring Solvation Interactions of Deep Eutectic Solvents Formed by Metal Chlorides and an Imidazolium Salt by Inverse Gas Chromatography

Deep eutectic solvents (DESs) represent a class of solvents that offer a number of advantages including minimal toxicity, affordability, low vapor pressure, and simple, environmentally friendly preparation methods. They have found utility in areas, such as gas absorption, metal plating, and extractions. However, the relationship between their solvation properties and chemical composition remains poorly understood. In this study, a broad range of Type I DESs composed of metal chlorides and an imidazolium salt were prepared, employed as gas chromatographic stationary phases, and characterized using the Abraham solvation parameter model by inverse gas chromatography. The Abraham solvation parameter model allows for the study of DES solvation properties and the effects of varying structural components on system constants using a linear-free energy relationship. The DESs were investigated by systematically varying their composition, including the type of metal chloride and the molar ratio between the metal chloride and the imidazolium salt. The results show that hydrogen bond acidity, hydrogen bond basicity, and dipolarity/polarizability interactions are strongly influenced by the type of metal chloride within the DES and the ratio of metal chloride to imidazolium salt in the eutectic mixture. Furthermore, a column pretreatment procedure is presented that enables the effective coating of highly polar DESs onto open tubular capillary gas chromatography columns.

deep eutectic solvent↗

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science↗

Organo-mineral interactions in active layer and permafrost soils along aging Arctic landscapes

Rising temperatures are accelerating permafrost thaw, exposing large soil organic carbon (SOC) stocks to microbial decomposition with implications for global climate. Understanding how permafrost carbon is stored and protected through associations with minerals is critical for predicting its vulnerability to decomposition upon thaw. However, how landscape age, substrate chemistry, and soil depth influence mineral associations remain relatively unexplored. We investigated organo-mineral associations in active layer and permafrost soils across a landscape age and geochemical gradient on Alaska’s North Slope, spanning three glaciated (~11,500–125,000 years) and one unglaciated site. Using selective dissolution extractions, X-ray diffraction, and Mössbauer spectroscopy, we characterized minerals and their relationship with SOC. The three recently deglaciated sites had low soil pH that decreased with age and greater abundances of pyrophosphate- and oxalate-extractable Al and Fe, whereas the oldest unglaciated site exhibited near-neutral pH, greater pyrophosphate-extractable Ca, and distinct mineralogy. Across sites, SOC was positively associated with Al and Fe mineral phases, with stronger relationships in acidic soils. Pyrophosphate-extractable Ca also showed strong relationships with SOC at the acidic sites (up to ~10x greater), suggesting that Ca-mediated protection may operate beyond traditionally recognized high-pH soils. Permafrost soils showed depth-related changes in pH, SOC, and Fe mineralogy, suggesting chemically active, heterogeneous layers may shape mineral dynamics and associated carbon. Our results highlight how landscape age, parent material, and depth create distinct geochemical environments that govern mineral-organic associations. As thaw exposes soil to new conditions, these mineral-mediated protection mechanisms may be altered, potentially affecting the permafrost carbon-climate feedback.

Synthetic Biology↗

Enhancing dimensionality prediction in hybrid metal halides via feature engineering and class-imbalance mitigation

We present a machine learning (ML) framework for predicting the structural dimensionality of hybrid metal halides (HMHs), including organic-inorganic perovskites, using a combination of chemically-informed feature engineering and advanced class-imbalance handling techniques. This study is motivated by the small and highly imbalanced nature of experimentally available HMH datasets, which limits the applicability and reliability of conventional ML approaches. The dataset, consisting of 494 HMH structures, is highly imbalanced across dimensionality classes (0D, 1D, 2D, 3D), posing significant challenges to predictive modeling. To mitigate this limitation, the dataset was augmented to 1336 samples using the synthetic minority oversampling technique, enabling improved learning of underrepresented dimensionality classes while preserving chemically meaningful feature relationships. We developed interaction-based descriptors designed to capture coupled steric and polarity effects relevant to dimensionality prediction, which are not readily captured by standard single-parameter or composition-only descriptors. These descriptors are integrated into a multi-stage workflow combining feature selection, ensemble stacking, and performance optimization. Our approach significantly improves F1-scores for underrepresented classes, achieving robust cross-validation performance across all dimensionalities. This work demonstrates a generalizable strategy for extracting reliable and interpretable structure–dimensionality relationships from limited experimental data, enabling pre-synthesis screening of organic cations and providing a practical blueprint for small-data ML in hybrid materials systems.

36 MATERIALS SCIENCE↗

Comparative Performance Evaluation of Large Language Models for Extracting Molecular Interactions and Pathway Knowledge

Understanding the interactions and regulatory relationships among biomolecules is essential for deciphering complex biological systems and elucidating the mechanisms behind diverse biological functions. Traditionally, the collection of such molecular interaction data has relied on expert curation, a process that is both time-consuming and labor-intensive. To address these limitations, this study explores the use of large language models (LLMs) to automate the genome-scale extraction of molecular interaction knowledge. Here, we evaluate the performance of various LLMs on key biological tasks, including the identification of protein-protein interactions, detection of genes associated with pathways influenced by low-dose radiation, and inference of gene regulatory relationships. Our findings demonstrate that larger LLMs tend to perform better, particularly in extracting intricate gene and protein interactions. Despite their strengths, these models face challenges in recognizing functionally diverse gene groups and highly correlated regulatory relationships. Through a comprehensive analysis using established molecular interaction and pathway databases, we show that LLMs possess the potential to identify relevant biomolecules and predict their interactions, offering valuable insights and marking a significant step toward AI-driven biological knowledge discovery.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Link Scheduling in Satellite Networks via Machine Learning Over Riemannian Manifolds

Low Earth Orbit (LEO) satellites play a crucial role in enhancing global connectivity, serving a complementary solution to existing terrestrial systems. In wireless networks, scheduling is a vital process that allocates time-frequency resources to users for interference management. However, LEO satellite networks face significant challenges in scheduling their links towards ground users due to the satellites’ mobility and overlapping coverage. This paper addresses the dynamic link scheduling problem in LEO satellite networks by considering spatio-temporal correlations introduced by the satellites’ movements. The first step in the proposed solution involves modeling the network over Riemannian manifolds, thanks to their representation as symmetric positive definite matrices. We introduce two machine learning (ML)-based link scheduling techniques that model the dynamic evolution of satellite positions and link conditions over time and space. To accurately predict satellite link states, we present a recurrent neural network (RNN) over Riemannian manifolds, which captures spatio-temporal characteristics over time. Furthermore, we introduce a separate model, the convolutional neural network (CNN) over Riemannian manifolds, which captures geometric relationships between satellites and users by extracting spatial features from the network topology across all links. Simulation results demonstrate that both RNN and CNN over Riemannian manifolds deliver comparable performance to the fractional programming-based link scheduling (FPLinQ) benchmark. Remarkably, unlike other ML-based models that require extensive training data, both models only need 30 training samples to achieve over 99% of the sum rate while maintaining similar computational complexity relative to the benchmark.

42 ENGINEERING↗