Search NASA⌕ Search

SEARCH · Search NASA

Results for “relationship extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES↗

Antarctic Photochemistry: Uncertainty Analysis

Understanding the photochemistry of the Antarctic region is important for several reasons. Analysis of ice cores provides historical information on several species such as hydrogen peroxide and sulfur-bearing compounds. The former can potentially provide information on the history of oxidants in the troposphere and the latter may shed light on DMS-climate relationships. Extracting such information requires that we be able to model the photochemistry of the Antarctic troposphere and relate atmospheric concentrations to deposition rates and sequestration in the polar ice. This paper deals with one aspect of the uncertainty inherent in photochemical models of the high latitude troposphere: that arising from imprecision in the kinetic data used in the calculations. Such uncertainties in Antarctic models tend to be larger than those in models of mid to low latitude clean air. One reason is the lower temperatures which result in increased imprecision in kinetic data, assumed to be best characterized at 298K. Another is the inclusion of a DMS oxidation scheme in the present model. Many of the rates in this scheme are less precisely known than are rates in the standard chemistry used in many stratospheric and tropospheric models.

Stewart, Richard W.↗

From Surface Chlorophyll a to Phytoplankton Community Composition in Oceanic Waters

The objective of the present study is to examine the potential of using the near-surface total chlorophyll a concentration (C(sub surf)), as it can be derived from ocean color observation, to infer the column-integrated and the vertical distribution of the phytoplanktonic biomass, both in a quantitative way and in a qualitative way (z.e., in terms of community structure). Within this context, a large HPLC (High Performance Liquid Chromatography) pigment database has been analyzed. It includes 2419 vertical pigment profiles, all sampled in Case-1 waters with various trophic states. The relationshps between C(sub surf) and the total chlorophyll alpha vertical distribution, as previously derived by Morel and Berthon, are fully confirmed, as the present results coincide with the previous ones. This agreement allows to go further, namely to examine the possibility of extracting relationships between C(sub surf) and the vertical composition of the algal assemblages. Thanks to the detailed pigment composition available from HPLC measurements, the contribution of three size classes (micro-, nano-, and pico-phytoplankton) to the local total chlorophyll a concentration can be assessed. Corroborating previous findings (e.g., large species dominate in eutrophc environments, whereas tiny phytoplankton prevail in oligotrophic zones), the results lead to a statistically based parameterization. The predictive skill of this parameterization is successfully tested on a separate data set. With such a tool, the vertical total chlorophyll a profiles associated with each size class can be inferred from the sole knowledge of C(sub surf). By combining this tool with satellite ocean color data, it becomes conceivable to quantify on a global scale the phytoplankton biomass associated with each of the three size classes.

Uitz, Julia↗

Understanding Machine Learning in Earth Science: A Natural Language Processing Approach

Machine learning (ML) is being increasingly utilized in Earth science research. Benefits of ML include efficiency, reduction of human error, and ability to extract hidden patterns within data. However, the mutual lack of each other’s domain knowledge by ML and Earth science stands as a barrier to timely and effective implementation. Earth science, in particular, faces challenges in generating sample data, compared to those of traditional ML problems such as face recognition or stock predictions, where data is abundant and not lacking in ground truth, which is necessary for labeling. Earth science data are more varying in formats, such as HDF5 and image resolutions, and are not standardized across instruments, even within a given Earth science discipline. Previous studies have been done to outline the specific challenges that Earth science faces with ML, while others have focused on using existing publications to mine information efficiently. Other resources such as Scikit-Learn have developed decision trees for choosing appropriate machine learning algorithms, but application within Earth science subjects becomes much more complex. For the current study, we propose a methodology and tool that aids in implementation of ML in Earth science using natural language processing (NLP). Our work comprises three main parts: (1) analyzing existing publications related to ML and Earth science, using natural language processing: (2) extracting from the publications information on ML models subjects in Earth Science: and (3) visualizing the extracted relationships as a network graph. The resulting network graph should aid the Earth science communities in applying optimal ML algorithms and guiding data preparation through visualization of similar studies. The network graph and analysis of document similarity will be the basis of our next step, which is to develop a decision tree for selecting optimal machine learning methodologies for specified Earth science applications.

Zheng, Laura↗

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING↗

Automated Generation of Graph-based Cyber Threat Intel

With the advancement of AI technology and tools, specifically in the cybersecurity domain, both cyber defenders and threat actors are continuously adapting the use of these capabilities to expedite their operations. With this phenomenon, threat intelligence that is up to date, refreshable, and has relevant context to a specific threat becomes more and more important as it enables cybersecurity professionals to gain insight into relevant data and relationships to guide their operations. This project enables users to frequently aggregate threat intelligence from various sources, such as vendor vulnerability advisories affecting critical infrastructure, malware reports, and adversary writeups into a centralized, standardized database. The project utilizes the Structured Threat Intelligence eXpression (STIX) for a standardized, shareable threat intelligence data format and Neo4j as a graph database solution to store STIX nodes and relationships. Initial results of the project include datasets of over 8,000 nodes and 20,000 relationships extracted from over 500 data sources that have been released within the past month.

Threat Intelligence↗

MechBERT: Language Models for Extracting Chemical and Property Relationships about Mechanical Stress and Strain

Language models are transforming materials-aware naturallanguage processing by enabling the extraction of dynamic, context-rich information from unstructured text, thus, moving beyond the limitations of traditional information-extraction methods. Moreover, small language models are on the rise because some of them can perform better than large language models (LLMs) when given domain-specific questionanswer tasks, especially about an application area that relies on a highly specialized vernacular, such as materials science. We therefore present a new class of MechBERT language models for understanding mechanical stress and strain in materials. These employ Bidirectional Encoder Representations for transformer (BERT) architectures. We showcase four MechBERT models, all of which were pretrained on a corpus of documents that are textually rich in chemicals and their stress–strain properties and were fine-tuned on question-answering tasks. We evaluated the level of performance of our models on domain-specific as well as general English-language question-answer tasks and also explored the influence of the size and type of BERT architectures on model performance. We find that our MechBERT models outperform BERT-based models of the same size and maintain relevancy better than much larger BERT-based models when tasked with domain-specific question-answering tasks within the stress–strain engineering sector. These small language models also enable much faster processing and require a much smaller fraction of data to pretrain them, affording them greater operational efficiency and energy sustainability than LLMs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

metnet-direct-auxiliary

Code to reveal the relationship between fungal metabolomic outputs and the exogeneous treatments triggering their production. Two routes to extract the relationships: (1) direct route - for known and putative metabolites induced by treatments, and (2) for discovering unknown analytes induced by treatments. The code outputs the corresponding networks and also network-based metrics to rank the metabolomic outputs and treatments.

Meena, MuraliG↗

BNF C-band Scanning ARM Precipitation Radar 2nd Generation (CSAPR-2) Extracted Radar Columns and In-Situ Sensors (RadCLss)

Corrected Moments in Antenna Coordinates (CMAC) calculates quantitative precipitation estimates (QPE) from empirical relationships based on equivalent radar reflectivity factor, specific differ- ential phase, and specific attenuation. To evaluate these empirical relationships, the Extracted Radar Columns and In-Situ Sensors (RADclss) product was developed. Utilizing Py-ART, RAD- clss extracts CMAC radar columns above various ARM and partner locations. These columns are then spatiotemporally synced with in-situ observations at the surface utilizing the Atmospheric data Community Toolkit (ACT; Theisen et al. 2025), allowing direct comparison of radar parameters with rain gauges and laser disdrometers for further investigation.

bankhead↗

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT↗

From tides to seasons: How cyclic tidal drivers and plant physiology interact to affect carbon cycling at the terrestrial-estuarine boundary (Final technical report)

Coastal ecosystems are among the most biologically and biogeochemically active and diverse systems on Earth. Because they act as important linkages between terrestrial ecosystems and the open ocean, their incorporation in Earth system models (ESMs) is critical to predict coastal and global responses to environmental changes. However, they vary greatly in the magnitude of tides and the volume and timing of freshwater input from land, making it challenging to model the major biogeochemical reactions that control productivity and greenhouse gas emissions across coastal terrestrial aquatic interfaces (TAIs). Our overall objective was to improve mechanistic process understanding and modeling of tidal wetland hydro-biogeochemistry in coastal TAIs. We established a new flux tower site (Ameriflux US-PLo) in the oligohaline part of the Parker River to continuously monitor ecosystem-scale carbon fluxes under temporally varying salinity conditions. The site is co-located with long-term monitoring plots of the Plum Island Ecosystems LTER project. We installed wells and redox sensors in the marsh interior and creek bank, established biomass monitoring plots and deployed novel optode sensors in both locations. We used this data to parameterize plant-mediated transport in PFLOTRAN and tested the impact of soil heterogeneity on porewater constituents and gas fluxes. We collected observations of root oxygen release with a novel planar optode system in the field. Flux data collected during the measurement period encompasses a large variation in salinity ranging from drought to record precipitation years. We developed a method to extract functional relationships from the flux data using artificial neural networks, identifying salinity thresholds for CH 4 fluxes. Finally, we are using the coupled ELM-PFLOTRAN model to test the impact of antecedent hydrological conditions on the salinity-CH 4 flux relationship. This grant contributed to the professional development of one postdoc, three research assistants and one graduate student. The sensor data has been shared with external collaborators.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of aerodynamic derivatives from a magnetic balance system

The dynamic testing of a model in the University of Virginia cold magnetic balance wind-tunnel facility is expected to consist of measurements of the balance forces and moments, and the observation of the essentially six degree of freedom motion of the model. The aerodynamic derivatives of the model are to be evaluated from these observations. The basic feasibility of extracting aerodynamic information from the observation of a model which is executing transient, complex, multi-degree of freedom motion is demonstrated. It is considered significant that, though the problem treated here involves only linear aerodynamics, the methods used are capable of handling a very large class of aerodynamic nonlinearities. The basic considerations include the effect of noise in the data on the accuracy of the extracted information. Relationships between noise level and the accuracy of the evaluated aerodynamic derivatives are presented.

Raghunath, B. S.↗

Amplifying ribbon extensometer for measuring film and fabric strain

Variations of a low-cost amplifying linear threshold extensometer are presented in detail for high or low strain applications. Derivations of scale relationships and extraction forces are included with experimental correlations and analyses given on performance, attachment problems, gain selection, gage and base material compatibility, and zero setting techniques. Flight applications of unamplified gages on parawing deployment tests are noted. The amplified gages perform accurately in laboratory tests and further experience is needed on performance under dynamic environments.

Alley, V. L., Jr.↗

Investigation of the Performance and Explainability Tradeoffs for Machine-Learning Models for Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants

Predictive maintenance (PdM) has shown great potential for achieving substantial cost savings and enhancing the economic competitiveness of nuclear power plants (NPPs) in today's energy market. Among the different modeling approaches that exist, machine learning (ML) tools in particular have a demonstrated ability to handle high dimensional and multivariate data and to extract hidden relationships within data in industrial environments. While ML methods show great potential, their lack of explainability---especially for black-box models---is a major hurdle to their adoption. Moreover, considering the supposed trade-off between explainability and performance challenges, careful consideration must be made as to which of these quality aspects takes precedence in light of multiple modeling options, resource availability, and domain characteristics. The present work evaluates the performance of six ML models, each with a different degree of explainability, in classifying the conditions of circulating water pumps (CWPs) by utilizing sensor data from nuclear power plants. To determine the drivers behind the trade-offs presented by this array of models, this work also tests different combinations of CWP units as the training and testing data, degrees of data imbalance, and objective functions for hyperparameter tuning. It was found that black-box models tend to afford superior performance in cases where there are far more instances of one type of labeled data than of any other type. It is recommended that a guided procedure be followed for designing and delivering an ML system that is sufficiently explainable to all involved stakeholders.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Routing Algorithm Exploits Spatial Relations

A recently developed routing algorithm for broadcasting in an ad hoc wireless communication network takes account of, and exploits, the spatial relationships among the locations of nodes, in addition to transmission power levels and distances between the nodes. In contrast, most prior algorithms for discovering routes through ad hoc networks rely heavily on transmission power levels and utilize limited graph-topology techniques that do not involve consideration of the aforesaid spatial relationships. The present algorithm extracts the relevant spatial-relationship information by use of a construct denoted the relative-neighborhood graph (RNG).

Okino, Clayton↗

New Developments in the Embedded Statistical Coupling Method: Atomistic/Continuum Crack Propagation

A concurrent multiscale modeling methodology that embeds a molecular dynamics (MD) region within a finite element (FEM) domain has been enhanced. The concurrent MD-FEM coupling methodology uses statistical averaging of the deformation of the atomistic MD domain to provide interface displacement boundary conditions to the surrounding continuum FEM region, which, in turn, generates interface reaction forces that are applied as piecewise constant traction boundary conditions to the MD domain. The enhancement is based on the addition of molecular dynamics-based cohesive zone model (CZM) elements near the MD-FEM interface. The CZM elements are a continuum interpretation of the traction-displacement relationships taken from MD simulations using Cohesive Zone Volume Elements (CZVE). The addition of CZM elements to the concurrent MD-FEM analysis provides a consistent set of atomistically-based cohesive properties within the finite element region near the growing crack. Another set of CZVEs are then used to extract revised CZM relationships from the enhanced embedded statistical coupling method (ESCM) simulation of an edge crack under uniaxial loading.

Saether, E.↗