Search NASASearch

SEARCH · Search NASA

Results for “mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics

An RNA ligase partner for the prokaryotic protein-only RNase P: insights into the functional diversity of RNase P from genome mining

RNase P can use either an RNA- or a protein-based active site to catalyze 5'-maturation of transfer RNAs (tRNAs). This distinctive attribute in the biocatalytic repertoire raises questions about the underlying evolutionary driving forces, especially if each variant somehow affords a selective advantage under certain conditions. Upon mining all publicly available prokaryotic genomes and examining gene co-occurrence, we discovered that an RNA ligase with circularization activity was significantly overrepresented in genomes that contain the protein form of RNase P. This unexpected linkage inspires testable ideas to understand the bases for scenarios that might favor RNase P variants of different architectures/make-up.

HARP

An Exploratory Data Mining Investigation for Constructing a Publicly Sourced Dataset of Foreign Hypersonic Tests

This document details a data mining exercise that resulted in an exploratory dataset of publicly reported foreign (non-US) hypersonic vehicle test events. Using a combination of targeted English language searches and country-specific queries, the study aggregates information from digital news media, official press releases, and social media posts. The resulting list of events captures the publicly available accounts of foreign hypersonic tests, although it does not represent an exhaustive record. Limitations such as inconsistent reporting, translation challenges, and the inherently provisional nature of open-source data are acknowledged. This dataset serves as an initial reference point for further inquiries into high-speed atmospheric phenomena and may facilitate future efforts to correlate these events with geophysical measurements.

33 ADVANCED PROPULSION SYSTEMS

Recovery and Refining of Rare Earth Elements from Lignite Mine Wastes

The University of North Dakota (UND), in collaboration with a comprehensive team of technical, business and host-site partners, built on prior technology development to complete a front-end engineering and design (FEED) and business planning study to recover and refine rare earth elements (REE) and critical minerals (CM) from North Dakota (ND) lignite mine wastes. The end of project goal was to have an investment quality project and a committed team ready to commercialize the proposed technologies in a future construction and operations phase.

01 COAL, LIGNITE, AND PEAT

Energy-Efficient Selective Removal of Metal Ions from Mining Influenced Waters (MIW) Using H-Bonded Organic-Inorganic Framework (HOIFS) (CRADA Final Report)

The work developed a hydrogen-bonded organic–inorganic framework (HOIF), specifically zinc imidazole salicylaldoxime supramolecule (ZIOS), for selective Cu removal and recovery from acidic mining-impacted waters (AMD), with emphasis on scalable synthesis, membrane integration, durability in real AMD (RAMD), mechanistic understanding, and recovery/regeneration pathways. This research breaks new ground in selective resource recovery from AMD waters – no other sorbents can operate reliably in this pH range. This project is unique in what it has delivered and is of general use to the public for projects relating to critical materials recovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

"Mining" for Critical Minerals: Critical Minerals from Fossil Energy Waste Byproducts

NETL researcher Mengling Stuckman is an invited speaker for the panel discussion session of "Mining for Critical Minerals" at the Marcellus Shale Coalition event, "Shale 2.0: Learn the Facts about Upstream, Midstream and Downstream". Recent studies from DOE have shown that produced water and drill cuttings from the development of unconventional shale in Appalachia has a significant source of valuable critical minerals and will support bolstering America's supply chain security, putting Pennsylvania in a unique position to capitalize. The panel discussion facilitates obtaining an overview of pipeline capacity needs, downstream users for natural gas and effectively educating and engaging the public relevant to the Marcellus Shale community. The event also offers opportunities for local oil and gas industries, water and waste management companies to work with NETL and participate in FECM’s Critical Mineral program. Industrial feedbacks and participation as outcomes of this invited talk will accelerate the technology and knowledge transfer for the DOE’s Critical Mineral program and for DOE’s mission to unleash American Energy and lead in energy innovation.

critical minerals

Mining Thermophile Photosynthesis Genes: A Synthetic Operon Expressing Chloroflexota Species Reaction Center Genes in Rhodobacter sphaeroides

Photosynthesis is the foundation of the vast majority of life systems, and is therefore the most important bioenergetic process on earth. The greatest diversity of photosynthetic systems is found in microorganisms. However, our understanding of the biophysical and biochemical processes that transduce light into chemical energy is derived from a relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of predicted proteins that catalyze the initial photochemical reactions deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for the expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42–90 °C) springs available through the JGI IMG database, to generate a resource of diverse sequences that are potentially adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota

Voltage Mining for (De)lithiation-Stabilized Cathodes and a Machine Learning Model for Li-Ion Cathode Voltage

Advances in lithium-metal anodes have inspired interest in discovery of Li-free cathodes, most of which are natively found in their charged state. This is in contrast to today's commercial lithium-ion battery cathodes, which are more stable in their discharged state. In this study, we combine calculated cathode voltage information from both categories of cathode materials, covering 5577 and 2423 total unique structure pairs, respectively. The resulting voltage distributions with respect to the redox pairs and anion types for both classes of compounds emphasize design principles for high-voltage cathodes, which favor later Period 4 transition metals in their higher oxidation states and more electronegative anions like fluorine or polyanion groups. Generally, cathodes that are found in their charged, delithiated state are shown to exhibit voltages lower than those that are most stable in their lithiated state, in agreement with thermodynamic expectations. Deviations from this trend are found to originate from different anion distributions between redox pairs. In addition, a machine learning model for voltage prediction based on chemical formulas is trained and shows state-of-the-art performance when compared to two established composition-based ML models for material properties predictions, Roost and CrabNet.

25 ENERGY STORAGE

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science

Data Mining of Groundwater to Identify MAGs with Methane, Propane and Toluene Monooxygenases

Whole genome sequencing datasets, involving more than 600 groundwater samples, from nine countries, were analyzed to identify metagenome assembled genomes (MAGs) containing full operons for propane monooxygenase, soluble methane monooxygease, toluene monooxygenase and particulate ammonia/methane monooxygenase. The enzymes encoded by these genes are a focus of interest because of their ability to degrade common groundwater contaminants. Due to the large amount of data, sequence analyses involved more than 80 individual KBase narratives. The approach followed the KBase tutorial called "Metagenome-Assembled Genome Extraction from a Compost Microbiome Enrichment" The generated MAGs were exported from each individual narrative into separate summary KBase narratives for each monooxygenase. Three KBase narratives were generated for particulate ammonia/methane monooxygenase, due to the large number of MAGs identified.

59 BASIC BIOLOGICAL SCIENCES

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science

Membrane-based solvent extraction for the recovery of rare earths from phosphate mining process streams

This study reports on the capture of rare earth elements (REEs) from phosphate industry process streams, including phosphoric acid (PA) sludge and phosphogypsum (PG), using a membrane solvent extraction (MSX) process. While MSX has been proven effective for a relatively concentrated feed, its effectiveness for dilute REEs solutions remains unexplored. Investigated PA-sludge and PG particles contain total REEs concentrations of ∼1100 and ∼320 ppm, respectively. Acid leaching, implemented to dissolve the REEs, significantly dilutes the REEs concentration to ∼210 ppm for PA-sludge leachate and ∼60 ppm for PG leachate. These low concentrations, compounded by the higher levels of non-REE ions and radioactive species, uranium (U) and thorium (Th), poses challenges to the MSX process. Here, we demonstrated that N,N,N′,N′-tetraoctyl-diglycolamide (TODGA) selectively binds REEs from a >3 M nitric-acid leachate while effectively rejecting U and Th. Concentrations of light REEs in strip solution were doubled compared to the feed, while heavy REEs were preferentially extracted. Furthermore, >99% purity gypsum, free of U and Th, was precipitated during the acid leaching process, aiding separation by removing significant amounts of non-REEs species (e.g., calcium) prior to the MSX process. Molecular simulations support the experimental data, suggesting preferential separation of heavy over light REEs. Based on these results, a cost-effective integrated process including pretreatment, acid leaching, MSX, and wastewater treatment is proposed for the co-recovery of REEs, phosphoric acid, gypsum, and U. This study shows MSX as a technically and economically feasible process for the recovery of REEs from low-concentration process streams, offering advantages over conventional solvent extraction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH