Text Mining the Literature to Inform Experiments and Rationalize Impurity Phase Formation for BiFeO 3
Not Available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Scientific hypothesis generation represents a fundamental challenge in contemporary research due to exponentially expanding literature volumes and increasing disciplinary specialization. Large language models (LLMs) have emerged as transformative tools for automated scientific discovery, moving beyond traditional rule-based and literature-mining approaches. Four paradigmatic approaches define current LLM-driven hypothesis generation: direct prompting and fine-tuning methods, knowledge-enhanced frameworks integrating retrieval-augmented generation (RAG), multi-agent collaborative systems simulating research teams, and reasoning-focused approaches implementing cognitive architectures. Domain-specific applications demonstrate statistical equivalence to human expert performance in social psychology, experimental validation in biomedical research, and near-expert quality in astronomy. Evaluation methodologies encompass human expert assessment, LLM-as-judge frameworks, and comprehensive benchmarking systems. Technical challenges include hallucination management, knowledge integration limitations, and balancing novelty with feasibility. Future directions emphasize hybrid neural-symbolic architectures and sophisticated human-AI collaboration models for responsible scientific discovery acceleration.
Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.
A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.
Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.
The discovery and design of materials which can efficiently catalyze the oxygen reduction and evolution reactions at reduced temperatures is important for facilitating the widespread adoption of fuel cell and electrolyzer technologies. Numerous studies have produced correlations between catalytic properties, such as oxygen surface exchange or electrode area specific resistance (ASR), and properties of the catalyst material. However, correlations have historically been limited in scope (e.g., using only a few materials or at a single temperature) and it has been difficult to provide detailed assessments of their robustness. Here, in this study, we assess the ability of the O p-band center electronic structure descriptor, obtained from density functional theory (DFT) calculations, to correlate with oxygen surface exchange rates, diffusivities, and area specific resistances for a large database of perovskite oxide catalytic properties. By data mining the literature, we obtain 747 catalytic property value data points spanning 299 unique perovskite compositions from 313 studies. We assess linear correlations of each property with the O p-band center and find generally modest correlations that are qualitatively useful (prediction mean absolute errors of about 0.5 log units are typical), where the correlations are improved at higher temperatures (e.g., 800 °C vs. 500 °C) and significantly improve when considering fits to the subset of materials which have multiple independent measurements. These findings suggest that the spread of property data is significantly influenced by experimental uncertainty, and subsequent measurements of additional materials will likely improve the O p-band center correlations.
Not Available
Abstract Phase diagrams offer substantial predictive power for materials synthesis by identifying the stability regions of target phases. However, thermodynamic phase diagrams do not offer explicit information regarding the kinetic competitiveness of undesired by-product phases. Here we propose a quantitative and computable thermodynamic metric to identify synthesis conditions under which the propensity to form kinetically competing by-products is minimized. We hypothesize that thermodynamic competition is minimized when the difference in free energy between a target phase and the minimal energy of all other competing phases is maximized. We validate this hypothesis for aqueous materials synthesis through two empirical approaches: first, by analysing 331 aqueous synthesis recipes text-mined from the literature; and second, by systematic experimental synthesis of LiIn(IO 3 ) 4 and LiFePO 4 across a wide range of aqueous electrochemical conditions. Our results show that even for synthesis conditions that are within the stability region of a thermodynamic Pourbaix diagram, phase-pure synthesis occurs only when thermodynamic competition with undesired phases is minimized.
Two technology areas were advanced; a) novel molecules with improved performance in the end use application of organic corrosion inhibitors and flame retardant nylon polymers and b) development of a systematic process for identifying biomass-derived molecules with improved performance in end use applications. In total 17 novel organic corrosion inhibitors were identified that had significantly better performance than the commercial reference organic corrosion inhibitor and 7 novel nylons were synthesized with improved flame retardant properties relative to standard nylon-6,6. While an end-to-end systematic process for identifying biomass-derived molecules with improved end use performance was not completed, important progress was made computational tools for mining chemical structures from the literature and databases as well as establishing reaction network generation algorithms to aid in the discovery of novel molecules.
Mining multi-omics relationship networks from literature and databases
Surface crack detection and dimensional measurement at active mining sites present significant safety and operational challenges. Manual inspection methods are labor-intensive, spatially incomplete, and expose personnel to hazardous environments, while existing automated approaches have been developed primarily for concrete civil infrastructure and have not been validated on the complex, variable surfaces characteristic of mining environments. This dissertation presents an automated pipeline that integrates deep learning semantic segmentation with Structure-from-Motion photogrammetry to detect surface cracks and measure their aperture, length, and vertical displacement from standard RGB imagery acquired during routine Uncrewed Aerial Vehicle (UAV) survey operations, without requiring additional sensor hardware or manual measurement. The pipeline combines a U-Net architecture with an EfficientNet-B0 encoder, pretrained on the SDNET2018 concrete crack dataset and fine-tuned on a mining-specific dataset spanning laboratory concrete specimens, coal refuse impoundment embankments, and post-blast limestone quarry benches. Photogrammetric reconstruction is performed using COLMAP Structure-from-Motion and Multi-View Stereo, with crack segmentation masks projected into the reconstructed point cloud to enable three-dimensional vertical displacement measurement through local plane fitting and bimodal surface detection. The pipeline was validated across 36 controlled laboratory specimens at three imaging distances and four vertical displacement levels, achieving aperture measurement RMSE of 0.047 cm and R² of 0.954, and vertical displacement RMSE of 0.140 cm and R² of 0.966, against independent caliper measurements. Field application at a coal refuse impoundment in southwestern Pennsylvania detected 71 crack components across the embankment crest, with a dominant longitudinal crack exhibiting aperture values reaching 28 cm and a 95th percentile vertical displacement of 35.53 cm, consistent in magnitude and spatial distribution with simultaneously acquired LiDAR-derived estimates. Application across four post-blast limestone quarry bench datasets in California successfully characterized blast-induced fracture networks at ground sampling distances ranging from 0.59 to 1.23 cm/pixel, with detected crack geometries physically consistent with observable surface conditions at each site. The results demonstrate that deep learning-based crack detection and photogrammetric measurement can be integrated into routine UAV inspection workflows at mining sites, providing repeatable, scalable, and quantitative crack characterization across surface types, crack scales, and displacement magnitudes not previously addressed in the literature. The pipeline requires no dedicated surveying equipment beyond the UAV platforms already deployed at mine sites for survey and monitoring purposes, supporting practical adoption within existing operational workflows.
The demand for rare earth elements (REEs) has surged in recent years, driven by their crucial role in various industrial applications and their uneven geological distribution. As a result, urban mining from secondary resources, particularly coal and coal ash, has gained traction as a sustainable solution within a circular economy framework. This study highlights the significant presence of REEs in coal and coal ash, revealing that certain samples contain REE concentrations that rival traditional ores. Notably, coal ash has the potential to yield approximately 312,000 tons of REEs annually, far exceeding global demand. The research delves into advanced techniques for analyzing REEs, including elemental, isotopic, and mineralogical studies. Additionally, it explores innovative extraction methods such as the use of green solvents, nature-based solutions, and bioleaching and biosorption. By leveraging coal and its byproducts as secondary resources, this study underscores the opportunity to reduce dependence on conventional mining, enhancing the sustainability of REE recovery. Here, a comprehensive literature review was conducted to highlight technological advancements and emerging opportunities that can address current challenges in this field.
Gold nanoparticle synthesis recipes were extracted from the literature to obtain data-driven hypotheses for synthesis outcome morphology and size. Used images from https://Flaticon.com.
The Alliance of Genome Resources (Alliance) is an extensible coalition of knowledgebases focused on the genetics and genomics of intensively studied model organisms. The Alliance is organized as individual knowledge centers with strong connections to their research communities and a centralized software infrastructure, discussed here. Model organisms currently represented in the Alliance are budding yeast, Caenorhabditis elegans, Drosophila, zebrafish, frog, laboratory mouse, laboratory rat, and the Gene Ontology Consortium. The project is in a rapid development phase to harmonize knowledge, store it, analyze it, and present it to the community through a web portal, direct downloads, and application programming interfaces (APIs). Here, we focus on developments over the last 2 years. Specifically, we added and enhanced tools for browsing the genome (JBrowse), downloading sequences, mining complex data (AllianceMine), visualizing pathways, full-text searching of the literature (Textpresso), and sequence similarity searching (SequenceServer). We enhanced existing interactive data tables and added an interactive table of paralogs to complement our representation of orthology. To support individual model organism communities, we implemented species-specific “landing pages” and will add disease-specific portals soon; in addition, we support a common community forum implemented in Discourse software. We describe our progress toward a central persistent database to support curation, the data modeling that underpins harmonization, and progress toward a state-of-the-art literature curation system with integrated artificial intelligence and machine learning (AI/ML).
Controlled Environment Agriculture (CEA) offers high-yield, climate-resilient food production, but high energy and resource demands challenge its sustainability. This paper synthesizes technologies that can improve outcomes across six categories—energy, CO 2 utilization, building envelope, hardware, water, and process—plus colocation strategies. We evaluate 80 technologies and define ten implementation pathways bundling complementary technologies to reduce energy use, optimize water consumption, and minimize emissions. Regional application is demonstrated through five U.S. case studies spanning different climates. A logic framework guides pathway selection for case studies based on climate, infrastructure, and regulatory context, informing context-sensitive technology deployment. Results show energy intensity reductions of 3–55 %, ranging from energy management programs to comprehensive lighting retrofits; water savings of 20–40 % through closed-loop recirculation; and emissions reductions of 3–100 %, with strategic energy management achieving 3–5 % and renewable electricity paired with electrified heating achieving up to 100 %. Text mining revealed that energy, hardware, and process technologies account for 91 % of literature coverage. Water, building envelope, and CO 2 utilization remain underexplored, indicating priorities for future research. This integrative approach to technology assessment supports growers, developers, and policymakers in aligning CEA system design with local conditions, improving resource efficiency and addressing gaps in cross-domain technology coverage.
This report summarizes important nuances in local water concerns and potential climate impacts that could influence the roll-out of technologies associated with energy transitions. Current investments in clean energy technologies are very high, which is driving a lot of investments in related manufacturing (i.e., hydrogen, solar, wind, and batteries) and mining (e.g., lithium, copper, and graphite) around the world. To understand how water and climate dynamics could be influencing these activities, we conducted a phased literature review for three countries: China, Germany, and France. China was selected due to its global dominance in manufacturing of solar panels, batteries, and electrolyzers as well as production of rare earth elements while Germany and France were selected due to their emerging leadership in energy transitions-related manufacturing within the European Union. For each of these three nations, we identified areas where manufacturing is occurring within the country and then evaluated relevant water resources and climate impacts. Multiple sources were consulted for this review, including BloombergNEF, international reports, industry sources, peer-reviewed literature, climate data, and media coverage.
Wastewater is produced across nearly all human activities and requires treatment to safeguard human health and the natural environment. Treatment of wastewater often requires a large amount of thermal energy, resulting in wasted heat after the treatment process. Because hydrocarbon reforming needs both water and heat, the integration of wastewater treatment with hydrocarbon reforming, a process that produces synthesis gas rich in hydrogen, offers an excellent opportunity to utilize this waste heat and the impurities in wastewater to produce valuable hydrogen gas, to minimize waste from industrial processes, and to integrate water treatment with the hydrogen economy. Yet, no comprehensive literature review has been conducted to examine the integration of reforming and wastewater. To address this lack, we summarize the variety of catalytic reforming techniques available in the open literature and review the current literature on wastewater reforming with these techniques. Subsequently, we conduct a review of common types of wastewater contaminants and their possible effects on reforming catalyst performance and life. Lastly, three underexamined wastewater sources are identified, namely, oilfield wastewater, geothermal water, and mining and mineral processing wastewater, and their potential for future study as a reforming feedstock is examined.
Agent-based socio-technical modeling of medium- and heavy-duty (MDHD) electric vehicle (EV) adoption has the potential to provide analysis, prediction, and gui. This paper describes new applications of text analysis developed through machine learning (ML) to build and understand relevant topics and their saliency in the published discourse on adoption of MDHD EVs. This work contributes to the state of the art in topic mining models by defining a new metric of topic ranking (START) that quantifies the importance of predefined topics within the corpus using weighted results for predefined topics from two topic modeling approaches: Latent Dirichlet Allocation (LDA) and BERTopic. The START metric is then demonstrated in practice to model how academia and industry view the EV adoption process based on the respective texts published by these groups. Results show that academic literature places more emphasis on categories of interests such as norms/attitudes and adopter knowledge, while trade journals tend to emphasize long-term cost more than academia. The two bodies of literature agree on the importance of policy and incentives in MDHD EV adoption. Together these results illustrate the potential to use ML-based text analysis to populate the characteristics of agent-based socio-technical models.