Search NASA⌕ Search

SEARCH · Search NASA

Results for “data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Physics-Informed Neural Network (PINN) Prediction of Mixed Mass-Heat-Crystallization Limited Methane Hydrate Formation and Dissociation in Micro-Confinement

The creation and use of Physics-Informed Neural Networks (PINNs) for simulating the dynamics of methane hydrate formation and dissociation will be presented. The PINN framework's main benefit is its capacity to impose physical consistency with only a partial comprehension of the governing equations. This makes the algorithm especially useful for systems with little experimental evidence or a lack of theoretical knowledge. A strong basis for forecasting methane hydrate behavior over the verified operating ranges of 30.0-80.9 bar pressure and 1.0-4.0 K sub-cooling conditions is provided by the combination of conductive heat transfer equations and mixed mass-transfer–crystallization kinetics. PINNs were more accurate at predicting the mixed mass-heat-crystallization limited kinetics than conventional Artificial Neural Networks (ANNs), demonstrating remarkable predictive accuracy for methane hydrate production over the ANN model. The efficiency of incorporating physical limitations from first principles into machine learning frameworks for methane hydrate crystallizations is reinforced by these findings. For hydrate-related applications in energy generation, carbon sequestration, and climate modelling, our study establishes PINNs as a computational tool that is both scalable and efficient. The proven capacity to close the gap between conventional physics-based simulations and solely data-driven models creates new opportunities for expedited hydrate research and practical applications.

Hartman, Ryan L [NYU Tandon School of Engineering]↗

Mining Thermophile Photosynthesis Genes: A Synthetic Operon Expressing Chloroflexota Species Reaction Center Genes in Rhodobacter sphaeroides

Photosynthesis is the foundation of the vast majority of life systems, and is therefore the most important bioenergetic process on earth. The greatest diversity of photosynthetic systems is found in microorganisms. However, our understanding of the biophysical and biochemical processes that transduce light into chemical energy is derived from a relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of predicted proteins that catalyze the initial photochemical reactions deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for the expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42–90 °C) springs available through the JGI IMG database, to generate a resource of diverse sequences that are potentially adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota↗

Carbon Utilization and Storage Partnership of the Western United States

This technical report documents research conducted under DOE Award No. DE-FE0031837 focused on evaluating the feasibility of carbon capture, utilization, and storage (CCUS) systems in the central and western United States. The project integrated geologic characterization, reservoir simulation, infrastructure modeling, and economic analysis to assess CO₂ storage potential near industrial sources and develop strategies for transport and sequestration. The work included subsurface modeling, risk assessment, monitoring and verification (MRV) planning, and evaluation of regulatory pathways such as EPA Underground Injection Control (UIC) Class VI permitting and IRS 45Q tax credit eligibility. Results demonstrate the viability of multiple storage approaches, including saline formations, enhanced coalbed methane recovery, and basalt mineralization, supported by data-driven workflows and regional analyses. The project also produced permitting templates, technology transfer activities, and stakeholder engagement efforts to support deployment readiness. These findings contribute to the development of scalable, economically viable CCUS systems and provide a repeatable framework for future carbon management projects.

20 FOSSIL-FUELED POWER PLANTS↗

Image-Based Fracture Surface Defect Characterization Methods for Additively Manufactured Ti-6Al-4V Tested in Fatigue

Abstract Fatigue initiation in additively manufactured samples/parts often occurs at processed-induced defects such as lack-of-fusion (LoF), keyhole, or other morphological/microstructural defects that have unique characteristics and measurable qualities. Attempts at identifying and minimizing such defects have utilized optimized processing conditions along with in situ and ex situ characterization that includes metallography and/or X-ray computed tomography (XCT). This paper highlights the benefits of using fracture surface analyses to detect and quantify defects that may not be detected by metallography/XCT due to sectioning and resolution limits. In addition to using manual quantification of fatigue initiating LoF and keyhole defects on fracture surfaces, image-based machine learning using convolutional neural networks such as U-Net were also used to automate the process. Statistical analyses were used to identify the extreme cases of defects that initiated and accelerated fatigue and to model the distribution of defect size and shape characteristics to distinguish the type of defect. Initial results show agreement between trained machine learning models and ground truth data in defect segmentation, and the distributions of defect characteristics are distinguishable to particular process-induced defect types.

Materials Science↗

Primary material supply configurations and domestic recycling for cost-effective battery material production in the US

Battery cathode active material costs hinge on regionally concentrated, price-volatile metal supply. Here. we construct a regional facility-level cost model based on over 80 global lithium, cobalt, and nickel mines, refineries, and battery-grade material plants. Our model yields aggregated lithium, nickel, manganese, and cobalt production material costs from 392 region-based supply configurations for five different cathode active materials. Focusing on the United States, all-domestic supply is 9–34% costlier than global average, increasing by cobalt content, while these shortfalls can be overcome by selective low-cost material imports. Furthermore, we analyze costs of two U.S.-based recycling facilities from primary data and techno-economic modelling and compare resulting cathode active material-level costs to primary supply. Although it is still significantly higher on cathode active material cost-level, rising end-of-life flows and lowered black-mass prices will, however, make secondary supply cost-competitive to domestic and foreign primary supply cost floors. Facility-level benchmarks reveal targeted import, scaling, and production cost optimization as levers for a resilient, cost-effective U.S. battery-material supply chain.

Energy↗

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗

The design and construction of the Chips water Cherenkov neutrino detector

Chips (CHerenkov detectors In mine PitS) was a prototype large-scale water Cherenkov detector located in northern Minnesota. The main aim of the R&D project was to demonstrate that construction costs of neutrino oscillation detectors could be reduced by at least an order of magnitude compared to other equivalent experiments. This article presents design features of the Chips detector along with details of the implementation and deployment of the prototype. While issues during and after the deployment of the detector prevented data taking, a number of key concepts and designs were successfully demonstrated.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Combining Machine Learning and Comparative Effectiveness Methodology to Study Primary Care Pharmacotherapy Pathways for Veterans With Depression

Our objective is to demonstrate an innovative method combining machine learning with comparative effectiveness research techniques and to investigate a hitherto unstudied question about the effectiveness of common prescribing patterns. For Operation Enduring Freedom/Operation Iraqi Freedom veterans with major depressive disorder, we generate pharmacotherapy pathways (of antidepressants) using process mining and machine learning. We select the medication episodes that were started at subtherapeutic doses by the first assigned primary care physician and observe the paths that those medication episodes follow. Using 2-stage least squares, we test the effectiveness of starting at a low dose and staying low for longer versus ramping up fast while balancing observable and unobservable characteristics of patients and providers through instrumental variables. We leverage predetermined provider practice patterns as instruments. We collected outpatient pharmacy data for selective serotonin reuptake inhibitors and selective norepinephrine reuptake inhibitors, patient and provider characteristics (as control variables), and the instruments for our cohort. All data were extracted for the period between 2006 and 2020. There is a statistically significant positive effect (0.68, 95% CI 0.11–1.25) of “ramping up fast” on engagement in care. When we examine the effect of “ramping up slow”, we see an insignificant negative impact on engagement in care (−0.82, 95% CI −1.89 to 0.25). As expected, the probability of drop-out also seems to have a negative effect on engagement in care (−0.39, 95% CI −0.94 to 0.17). We further validate these results by testing with medication possession ratios calculated periodically as an alternative engagement in care metric. Our findings contradict the “Start low, go slow” adage, indicating that ramping up the dose of an antidepressant faster has a significantly positive effect on engagement in care for our population.

60 APPLIED LIFE SCIENCES↗

Signal processing and spectral modeling for the BeEST experiment

The Beryllium Electron capture in Superconducting Tunnel junctions (BeEST) experiment searches for evidence of heavy neutrino mass eigenstates in the nuclear electron capture decay of 7 Be by precisely measuring the recoil energy of the 7 Li daughter. In Phase III, the BeEST experiment has been scaled from a singl superconducting tunnel junction (STJ) sensor to a 36-pixel array to increase sensitivity and mitigate gamma-induced backgrounds. Phase III also uses a new continuous data acquisition system that greatly increases the flexibility for signal processing and data cleaning. Here, we have developed procedures for signal processing and spectral fitting that are sufficiently robust to be automated for large datasets. Furthermore, this article presents the optimized procedures before unblinding the majority of the Phase III dataset to search for physics beyond the standard model.

6 ≤ A ≤ 19↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

Desert-Urban System Integrated Atmospheric Monsoon (DUSTIEAIM) in the Southwestern United States Science Plan

The Desert-Urban System Integrated Atmospheric Monsoon (DUSTIEAIM) campaign is a groundbreaking, high-impact scientific mission that will transform how we understand and respond to energy and water challenges in one of America’s fastest-growing and most heat-stressed urban regions: Phoenix, Arizona. Starting in April 2026, this 18-month field campaign harnesses the full power of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility and an interdisciplinary science team including national laboratories, universities, and agencies with a broad range of subject-matter expertise. With cutting-edge instruments, active and passive ground-based sensors, radars, and integrated modeling, DUSTIEAIM will deliver the most comprehensive environmental data set ever collected for a desert-urban-agricultural interface.

54 ENVIRONMENTAL SCIENCES↗