Search NASA⌕ Search

SEARCH · Search NASA

Results for “database mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT↗

LIB Design Module for Grid Energy System Application

We will employ a machine learning approach with intelligent data mining and database construction to analyze enormous data repositories for identifying and extracting geographic-dependent cell design specifications from publicly accessible grid-scale energy storage usage databases in an automated way at scale.

Liu, Dianying [Pacific Northwest National Laborato↗

Descriptors of water aggregation

For this work, we rely on a total of 23 (cluster size, 8 structural, and 14 connectivity) descriptors to investigate structural patterns and connectivity motifs associated with water cluster aggregation. In addition to the cluster size n (number of molecules), the 8 structural descriptors can be further categorized into (i) one-body (intramolecular): covalent OH bond length (r OH ) and HOH bond angle (θ HOH ), (ii) two-body: OO distance (r OO ), OHO angle (θ OHO ), and HOOX dihedral angle ($\phi$ HOOX ), where X lies on the bisector of the HOH angle, (iii) three-body: OOO angle (θ OOO ), and (iv) many-body: modified tetrahedral order parameter (q) to account for two-, three-, four-, five-coordinated molecules (q m , m = 2, 3, 4, 5) and radius of gyration (R g ). The 14 connectivity descriptors are all many-body in nature and consist of the AD, AAD, ADD, AADD, AAAD, AAADD adjacencies [number of hydrogen bonds accepted (A) and donated (D) by each water molecule], Wiener index, Average Shortest Path Length, hydrogen bond saturation (% HB), and number of non-short-circuited three-membered cycles, four-membered cycles, five-membered cycles, six-membered cycles, and seven-membered cycles. We mined a previously reported database of 4 948 959 water cluster minima for (H 2 O) n , n = 3–25 to analyze the evolution and correlation of these descriptors for the clusters within 5 kcal/mol of the putative minima. It was found that r OH and % HB correlated strongly with cluster size n, which was identified as the strongest predictor of energetic stability. Marked changes in the adjacencies and cycle count were observed, lending insight into changes in the hydrogen bond network upon aggregation. A Principal Component Analysis (PCA) was employed to identify descriptor dependencies and group clusters into specific structural patterns across different cluster sizes. The results of this study inform our understanding of how water clusters evolve in size and what appropriate descriptors of their structural and connectivity patterns are with respect to system size, stability, and similarity. The approach described in this study is general and can be easily extended to other hydrogen-bonded systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unifying Quantum Materials Modeling and Experiments: The Role of Machine Learning Interatomic Potentials

Computational experiments have emerged as a powerful complement to traditional experiments in the design of new materials. The development of machine learning (ML) and deep learning techniques, combined with database construction and data mining, has significantly enhanced traditional quantum mechanical methods. This synergy enables the rapid development of structure-property relationships. In this talk, I will discuss our recent efforts in applying Machine Learning Interatomic Potentials (MLIAPs) to accelerate materials modeling across various material classes and challenging applications where traditional methods fall short. First, I will highlight the success of MLIAPs in accurately modeling the melting behavior of complex materials. Our results demonstrate high fidelity with experimental observations and also with calculated reference melting temperatures. In the second application, I will discuss how MLIAPs are trained and applied to elucidate the interplay between segregation tendencies and surface reconstructions in CuNi alloys under oxidizing conditions. A key factor in the success of these MLIAP applications is the design of minimalistic yet flexible datasets along with a computational framework for training MLIAPs.

Saidi, Wissam↗

A database of thermally activated delayed fluorescent molecules auto-generated from scientific literature with ChemDataExtractor

A database of thermally activated delayed fluorescent (TADF) molecules was automatically generated from the scientific literature. It consists of 25,482 data records with an overall precision of 82%. Among these, 5,349 records have chemical names in the form of SMILES strings which are represented with 91% accuracy; these are grouped in a subsidiary database. Each data record contains one of the following four properties: maximum emission wavelength (λ EM ), photoluminescence quantum yield (PLQY), singlet-triplet energy splitting (ΔE ST ), and delayed lifetime (τ D ). The databases were created through text mining using ChemDataExtractor, a chemistry-aware natural-language-processing toolkit, which has been adapted for TADF research. The text-mined corpus consisted of 2,733 papers from the Royal Society of Chemistry and Elsevier. To the best of our knowledge, these databases are the first databases that have been auto-generated for TADF molecules from existing publications. The databases have been publicly released for experimental and computational applications in the TADF research field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A systematic analysis and data mining of opioid-related adverse events submitted to the FAERS database

The opioid epidemic has become a serious national crisis in the United States. An indepth systematic analysis of opioid-related adverse events (AEs) can clarify the risks presented by opioid exposure, as well as the individual risk profiles of specific opioid drugs and the potential relationships among the opioids. In this study, 92 opioids were identified from the list of all Food and Drug Administration (FDA)-approved drugs, annotated by RxNorm and were classified into 13 opioid groups: buprenorphine, codeine, dihydrocodeine, fentanyl, hydrocodone, hydromorphone, meperidine, methadone, morphine, oxycodone, oxymorphone, tapentadol, and tramadol. A total of 14,970,399 AE reports were retrieved and downloaded from the FDA Adverse Events Reporting System (FAERS) from 2004, Quarter 1 to 2020, Quarter 3. After data processing, Empirical Bayes Geometric Mean (EBGM) was then applied which identified 3317 pairs of potential risk signals within the 13 opioid groups. Based on these potential safety signals, a comparative analysis was pursued to provide a global overview of opioid-related AEs for all 13 groups of FDA-approved prescription opioids. The top 10 most reported AEs for each opioid class were then presented. Both network analysis and hierarchical clustering analysis were conducted to further explore the relationship between opioids. Results from the network analysis revealed a close association among fentanyl, oxycodone, hydrocodone, and hydromorphone, which shared more than 22 AEs. In addition, much less commonly reported AEs were shared among dihydrocodeine, meperidine, oxymorphone, and tapentadol. On the contrary, the hierarchical clustering analysis further categorized the 13 opioid classes into two groups by comparing the full profiles of presence/absence of AEs. The results of network analysis and hierarchical clustering analysis were not only consistent and cross-validated each other but also provided a better and deeper understanding of the associations and relationships between the 13 opioid groups with respect to their adverse effect profiles.

Research & Experimental Medicine↗

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES↗

Computer-aided, resistance gene-guided genome mining for proteasome and HMG-CoA reductase inhibitors

Secondary metabolites (SMs) are biologically active small molecules, many of which are medically valuable. Fungal genomes contain vast numbers of SM biosynthetic gene clusters (BGCs) with unknown products, suggesting that huge numbers of valuable SMs remain to be discovered. It is challenging, however, to identify SM BGCs, among the millions present in fungi, that produce useful compounds. One solution is resistance gene-guided genome mining, which takes advantage of the fact that some BGCs contain a gene encoding a resistant version of the protein targeted by the compound produced by the BGC. The bioinformatic signature of such BGCs is that they contain an allele of an essential gene with no SM biosynthetic function, and there is a second allele elsewhere in the genome. We have developed a computer-assisted approach to resistance gene-guided genome mining that allows users to query large databases for BGCs that putatively make compounds that have targets of therapeutic interest. Working with the MycoCosm genome database, we have applied this approach to look for SM BGCs that target the proteasome β6 subunit, the target of the proteasome inhibitor fellutamide B, or HMG-CoA reductase, the target of cholesterol reducing therapeutics such as lovastatin. Our approach proved effective, finding known fellutamide and lovastatin BGCs as well as fellutamide- and lovastatin-related BGCs with variations in the SM genes that suggest they may produce structural variants of fellutamides and lovastatin. Gratifyingly, we also found BGCs that are not closely related to lovastatin BGCs but putatively produce novel HMG-CoA reductase inhibitors.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying Performance Advantaged Biobased Chemicals Utilizing Bioprivileged Molecules

Two technology areas were advanced; a) novel molecules with improved performance in the end use application of organic corrosion inhibitors and flame retardant nylon polymers and b) development of a systematic process for identifying biomass-derived molecules with improved performance in end use applications. In total 17 novel organic corrosion inhibitors were identified that had significantly better performance than the commercial reference organic corrosion inhibitor and 7 novel nylons were synthesized with improved flame retardant properties relative to standard nylon-6,6. While an end-to-end systematic process for identifying biomass-derived molecules with improved end use performance was not completed, important progress was made computational tools for mining chemical structures from the literature and databases as well as establishing reaction network generation algorithms to aid in the discovery of novel molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Critical Assessment of Electronic Structure Descriptors for Predicting Perovskite Catalytic Properties

The discovery and design of materials which can efficiently catalyze the oxygen reduction and evolution reactions at reduced temperatures is important for facilitating the widespread adoption of fuel cell and electrolyzer technologies. Numerous studies have produced correlations between catalytic properties, such as oxygen surface exchange or electrode area specific resistance (ASR), and properties of the catalyst material. However, correlations have historically been limited in scope (e.g., using only a few materials or at a single temperature) and it has been difficult to provide detailed assessments of their robustness. Here, in this study, we assess the ability of the O p-band center electronic structure descriptor, obtained from density functional theory (DFT) calculations, to correlate with oxygen surface exchange rates, diffusivities, and area specific resistances for a large database of perovskite oxide catalytic properties. By data mining the literature, we obtain 747 catalytic property value data points spanning 299 unique perovskite compositions from 313 studies. We assess linear correlations of each property with the O p-band center and find generally modest correlations that are qualitatively useful (prediction mean absolute errors of about 0.5 log units are typical), where the correlations are improved at higher temperatures (e.g., 800 °C vs. 500 °C) and significantly improve when considering fits to the subset of materials which have multiple independent measurements. These findings suggest that the spread of property data is significantly influenced by experimental uncertainty, and subsequent measurements of additional materials will likely improve the O p-band center correlations.

30 DIRECT ENERGY CONVERSION↗

Updates to the Alliance of Genome Resources central infrastructure

The Alliance of Genome Resources (Alliance) is an extensible coalition of knowledgebases focused on the genetics and genomics of intensively studied model organisms. The Alliance is organized as individual knowledge centers with strong connections to their research communities and a centralized software infrastructure, discussed here. Model organisms currently represented in the Alliance are budding yeast, Caenorhabditis elegans, Drosophila, zebrafish, frog, laboratory mouse, laboratory rat, and the Gene Ontology Consortium. The project is in a rapid development phase to harmonize knowledge, store it, analyze it, and present it to the community through a web portal, direct downloads, and application programming interfaces (APIs). Here, we focus on developments over the last 2 years. Specifically, we added and enhanced tools for browsing the genome (JBrowse), downloading sequences, mining complex data (AllianceMine), visualizing pathways, full-text searching of the literature (Textpresso), and sequence similarity searching (SequenceServer). We enhanced existing interactive data tables and added an interactive table of paralogs to complement our representation of orthology. To support individual model organism communities, we implemented species-specific “landing pages” and will add disease-specific portals soon; in addition, we support a common community forum implemented in Discourse software. We describe our progress toward a central persistent database to support curation, the data modeling that underpins harmonization, and progress toward a state-of-the-art literature curation system with integrated artificial intelligence and machine learning (AI/ML).

59 BASIC BIOLOGICAL SCIENCES↗

An automated integrated web-based smart tool for open stope design

The Stability Graph is a widely used tool for the design of open stopes in underground mining. Many users of the Stability Graph still apply this design method manually. Although the manual approach has benefits, using multiple graphs and stability number computation charts for each stope surface is time-consuming, even for the experienced mining engineer. Current practice in the use of the method also limits data sharing. This paper presents a StopeSoft web-based tool for open stope stability prediction that is developed on the basis of the Stability Graph method and is available at openstope.com. StopeSoft incorporates flexibility in terms of Stability Graph options and incorporates additional critical factors often overlooked. As a web-based tool, StopeSoft encourages and makes data sharing possible globally, focused on expanding the database and improving the current limitations of the Stability Graph to provide practical, reliable solutions for mining engineers, consultants, and academics. The StopeSoft automated process facilitates the process of open stope stability prediction, saving time and minimizing potential human errors. Statistical treatment of the data accounts for the variability of input parameters to emphasize the probabilistic nature of the Stability Graph method. The probabilistic interpretation of the stability states of stope surfaces eliminates the false feeling of absolute stope performance based on its location on the Stability Graph , as implied by the deterministic approach.

58 GEOSCIENCES↗

Mining Thermophile Photosynthesis Genes: A Synthetic Operon Expressing Chloroflexota Species Reaction Center Genes in Rhodobacter sphaeroides

Photosynthesis is the foundation of the vast majority of life systems, and is therefore the most important bioenergetic process on earth. The greatest diversity of photosynthetic systems is found in microorganisms. However, our understanding of the biophysical and biochemical processes that transduce light into chemical energy is derived from a relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of predicted proteins that catalyze the initial photochemical reactions deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for the expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42–90 °C) springs available through the JGI IMG database, to generate a resource of diverse sequences that are potentially adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota↗

U.S. Freight Transload Facilities Dataset

The U.S. Freight Transload Facilities Dataset provides location information (latitude, longitude, zip, city, county, state)for more than 9,000 facilities across 50 U.S. States where freight may be transferred between waterways, railways, and roadways. The dataset lists the known modes and available direction(s) for freight transfers at each facility as of 2024. The U.S. Freight Transload Facilities dataset was built by mining and fusing several public sources, such as the USACE Master Docks Plus, the USDOT National Transportation Atlas Database (NTAD), files from the Intermodal Association of North America (IANA), and the industry publication Bulk Transloader. The dataset constitutes a key piece of a multimodal freight transportation network and routing algorithm developed by USACE-ERDC. The dataset is shared as a .csv file. The dataset is published for research purposes and should not be considered exhaustive or authoritative.

Peterson, Steven [ORNL] (ORCID:0000000287672998)↗

Detection and Association of Operational Events using DAS and Seismometers (FY 2025 Mid-Year Report)

This mid-year report summarizes ongoing work to identify anomalous vibration signals indicative of potential containment breaches. This work includes compiling continuous seismic datasets and testing and refining underground detection and geolocation techniques. In the first two quarters of FY25, we have completed two project work plan tasks: (1) creating a database of continuous waveforms and ground truth event data from multiple modalities and (2) refining and implementing a detection and association algorithm to create a catalog of anomalous underground activities. This report contains a summary of the seismic database including the continuous seismic data collected by a dense array of surface seismic stations above Pleasant Gap Mine, and continuous seismic data collected using subsurface distributed acoustic sensing (DAS) in the subsurface at Sanford Underground Research Facility (SURF) and the ground truth information gathered from both sites. This report also includes results from refining and applying a dynamic power spectral density detector to both continuous seismic datasets. Finally, the report provides an initial catalog of subsurface operational events from both sensing modalities.

58 GEOSCIENCES↗

NEWTS Integrated Dataset (version 1.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v1.0 provides water researchers, community leaders, and regulators with a unified and standardized energy-related wastewater stream database. This resource is derived from 27 state and federal entities, and scientific publications, and contains more than 400,000 sample records, many of which also provide geospatial information. The dataset includes data for several different energy-related wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, and power plant wastewater. The NEWTS Integrated Dataset was built to support environmentally prudent decision-making, explore treatment opportunities, and identify potential critical mineral sources. A subset of this novel resource is also featured on NETL NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

abandoned mine drainage↗