Search NASA⌕ Search

SEARCH · Search NASA

Results for “database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Annotation of DOM metabolomes with an ultrahigh resolution mass spectrometry molecular formula library

Current approaches to analyzing metabolomic data often rely on matching MS/MS fragmentation data to sparse libraries or databases. This approach results in limited identification of features, often with less than 10% of the dataset being annotated. A complementary approach is to assign molecular formula to features based on accurate mass measurements, but the platforms commonly used for metabolomics do not have the needed accuracy or resolving power to do this robustly, particularly for larger molecules. Using our newly modified analysis tool, CoreMS, we generated a library of molecular formula from pooled samples analyzed with LC-21T FT-ICR MS. This library successfully annotated approximately 53.2% of features identified from the exometabolome of marine diatom Phaeodactylum tricornutum – a nearly ten-fold increase over the 5.9% annotation rate achieved using a conventional MS/MS library matching approach. Using this FT-ICR MS library approach, we were able to differentiate differences in the exometabolome of P. tricornutum in iron replete and iron limited conditions, with 668 metabolites being differentially expressed (p < 0.05, 2 x intensity difference) under these conditions. The traditional MS/MS fragmentation-based annotation approach only annotated 61 of these metabolites, while our novel pipeline annotated 450 metabolites and revealed 12 metabolites that were significantly more abundant under low iron conditions. Our results demonstrate the utility of ultrahigh resolution mass spectrometry for generating more comprehensive and confident molecular annotations.

21T-FTICR-MS, CoreMS↗

Adsorbate-induced adatom formation on Au-Cu bimetallic alloys and its possible consequences for CO 2 electroreduction

The adsorbate-induced formation of sub-nanometer clusters on transition-metal single crystals observed in previous high-pressure microscopic studies hinted at the in-situ formation of unique active sites even on large nanoparticle catalysts. We propose that the adatom formation energy can be used as an energetic descriptor for the initial step toward the adsorbate-induced metal-cluster formation process. This descriptor can be efficiently computed using density functional theory (DFT) calculations and applied for screening and identification of metal catalysts where this phenomenon may play an important role in generating active sites in-situ. As a proof of concept, here, we construct an adatom formation energy database for three Au x Cu y alloys (x:y = 3:1, 1:1, or 1:3) and eighteen adsorbates (H, C, N, O, F, S, Cl, Br, I, CH x , NH x (x = 1 – 3), CO, NO, and OH) commonly involved in catalytic reactions. The energetics of adatom formation were examined in all cases where the (111) terrace, (211) step-edge, and (874) kink were the sources of the adatom. We demonstrate that the presence of an adsorbate could alter not only the energetics for adatom formation but also the elemental nature of the preferred adatom being formed. Using our database, we identified promising systems which favor adsorbate-induced adatom formation under near-ambient conditions. Specifically, CO-induced adatom formation on all three Au-Cu alloy surfaces could occur under CO 2 electroreduction (CO 2 RR) conditions. This phenomenon offers a qualitative explanation for the experimentally observed CO 2 RR activity on Au-Cu alloy catalysts. As a result, our methodology offers an easily expandable and efficient approach for large-scale catalyst screening with regards to adatom/cluster formation under reaction conditions and provides insight into the possible nature of active sites on alloy catalysts from a novel perspective.

Active site↗

Evaluation and optimization of flow boiling frictional pressure drop correlations using the data from traditional and next-generation refrigerants in a micro-fin tube

This study presents an experimental evaluation and optimization of flow boiling frictional pressure drop correlations for conventional and next-generation refrigerants in a horizontal micro-fin tube, with particular emphasis on the newly emerging refrigerant blends R-454C and R-455A, for which pressure-drop data in enhanced tubes remain limited. Experiments were conducted with R-410A, R-454C, R-455A, R-134a, R-1234yf, and R-1234ze(E) in a copper micro-fin tube with an inner diameter of 8.468 mm, over mass fluxes ranging from 100 to 300 kg/(m²·s) depending on the refrigerant, and evaporation temperatures of 7, 12, and 14 °C. Frictional pressure gradients were determined from measured total pressure drops after subtracting acceleration pressure drop, and the resulting database was used to assess four existing models: Kuo and Wang (1996), Cavallini et al. (1997), Goto et al. (2001), and Diani et al. (2014). The measured frictional pressure gradient increased with vapor quality and mass flux for all refrigerants and increased further at lower evaporation temperatures, with the overall trend strongly related to liquid viscosity. Among the four correlations, the Goto et al. (2001) model provided the best overall agreement with the measured data before optimization. To further improve prediction accuracy, the Kuo and Wang (1996) and Goto et al. (2001) models were optimized using the complete experimental database. After optimization, both models reduced the overall mean absolute deviation to below 15%, while the optimized Goto et al. (2001) model maintained the best and most consistent overall performance. The results provide new pressure-drop data for next-generation refrigerants and demonstrate that parameter optimization can significantly enhance the applicability of existing micro-fin-tube correlations.

Hu, Yifeng [ORNL] (ORCID:0000000242875185)↗

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES↗

Structure–Composition Relationships for Mg–Ni and Mg–Fe Olivine

Olivine is a dynamic and important mineral in the crust and mantle with relevance to processes important to climate change technology, such as geologic carbon storage and critical mineral recovery. In this work, we critically evaluated and compiled a new database of olivine diffraction data, lattice parameters, and composition to enable rapid Ni-Mg-Fe olivine composition determination. A compilation of olivine X-ray diffraction data and chemical compositions from both the literature and the International Centre for Diffraction Data (ICDD) powder database was assembled to plot both the forsterite-fayalite and forsterite-liebenbergite solid solution lines. Here we present an expanded dataset to delineate equations and relationships used for quantifying the correlations between olivine lattice parameters and chemical compositions in Mg 2 SiO 4 -Fe 2 SiO 4 (forsterite-fayalite) and Mg 2 SiO 4 -Ni 2 SiO 4 (forsterite-liebenbergite) olivine solid solution series.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Life Cycle Inventory Availability: Status and Prospects for Leveraging New Technologies

The demand for life cycle assessments (LCA) is growing rapidly, which leads to an increasing demand of life cycle inventory (LCI) data. While the LCA community has made significant progress in developing LCI databases for diverse applications, challenges still need to be addressed. This perspective summarizes the current data gaps, transparency, and uncertainty aspects of existing LCI databases. Additionally, we survey and discuss novel techniques for LCI data generation, dissemination, and validation. We propose key future directions for LCI development efforts to address these challenges, including leveraging scientific and technical advances such as the Internet of Things (IoT), machine learning, and blockchain/cloud platforms. Adopting these advanced technologies can significantly improve the quality and accessibility of LCI data, thereby facilitating more accurate and reliable LCA studies.

blockchain platforms↗

Remote Sensing Improves Multi‐Hazard Flooding and Extreme Heat Detection by Fivefold Over Current Estimates

The co‐occurrence of multiple hazards is of growing concern globally as the frequency and magnitude of extreme climate events increases. Despite studies examining the spatial distribution of such events, there has been little work in examining if all relevant life threatening and damaging hazards are captured in existing hazard databases and by common hazard metrics. For example, local/regional flash flooding events are seldom captured by optical satellite instruments and are subsequently excluded from global hazard databases. Similarly, the heat hazard definitions most frequently used in multi‐hazard studies inherently fail to capture events that are life‐threatening but climatologically within an expected range. Our goal is to determine the potential for increasing multi‐hazard event detection capabilities by inferring additional hazard footprints from widely accessible satellite data. We use daily precipitation and temperature satellite data to develop an open‐source framework that infers additional hazard footprints that are not included in traditional methods. With the state of Texas as our study area, we detected 2.5 times as many flood hazards, equivalent to $320 million in property and crop damages. Furthermore, our expanded heat hazard definition increases the impacted area by 56.6%, equivalent to 91.5 million km 2 over an 18 year period. Increasing hazard detection capabilities and expanding existing definitions of hazards using daily satellite data increases the temporal and spatial resolutions at which multi‐hazard events are detected. Having more complete data sets of all relevant hazard extents improves our ability to track global trends and more accurately determine the magnitude of hazard exposure inequities.

equity↗

RNA language models predict mutations that improve RNA function

Structured RNA lies at the heart of many central biological processes, from gene expression to catalysis. RNA structure prediction is not yet possible due to a lack of high-quality reference data associated with organismal phenotypes that could inform RNA function. We present GARNET (Gtdb Acquired RNa with Environmental Temperatures), a new database for RNA structural and functional analysis anchored to the Genome Taxonomy Database (GTDB). GARNET links RNA sequences to experimental and predicted optimal growth temperatures of GTDB reference organisms. Using GARNET, we develop sequence- and structure-aware RNA generative models, with overlapping triplet tokenization providing optimal encoding for a GPT-like model. Leveraging hyperthermophilic RNAs in GARNET and these RNA generative models, we identify mutations in ribosomal RNA that confer increased thermostability to the Escherichia coli ribosome. The GTDB-derived data and deep learning models presented here provide a foundation for understanding the connections between RNA sequence, structure, and function.

59 BASIC BIOLOGICAL SCIENCES↗

Metabolic interactions underpinning high methane fluxes across terrestrial freshwater wetlands

Current estimates of wetland contributions to the global methane budget carry high uncertainty, particularly in accurately predicting emissions from high methane-emitting wetlands. Microorganisms drive methane cycling, but little is known about their conservation across wetlands. To address this, we integrate 16S rRNA amplicon datasets, metagenomes, metatranscriptomes, and annual methane flux data across 9 wetlands, creating the Multi-Omics for Understanding Climate Change (MUCC) v2.0.0 database. This resource is used to link microbiome composition to function and methane emissions, focusing on methane-cycling microbes and the networks driving carbon decomposition. We identify eight methane-cycling genera shared across wetlands and show wetland-specific metabolic interactions in marshes, revealing low connections between methanogens and methanotrophs in high-emitting wetlands. Methanoregula emerged as a hub methanogen across networks and is a strong predictor of methane flux. In these wetlands it also displays the functional potential for methylotrophic methanogenesis, highlighting the importance of this pathway in these ecosystems. Collectively, our findings illuminate trends between microbial decomposition networks and methane flux while providing an extensive publicly available database to advance future wetland research.

54 ENVIRONMENTAL SCIENCES↗

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics↗

Accurate segmentation of localized corrosion in structural alloys via deep learning

This study presents a deep learning-based approach for the automated segmentation of corrosion damage in scanning electron microscopy (SEM) images. The proposed method enables rapid and accurate segmentation of corrosion features in these SEM images, making it highly suitable for real-time applications such as automated microscopy. Specifically, a dedicated corrosion segmentation database tailored for this task is constructed. The newly constructed dataset, alongside data from two public databases, are employed to jointly train a deep learning-based model modified with a texture refinement module. Compared to the same model without the texture refinement module, the refined model substantially enhances the efficacy and efficiency of corrosion segmentation. Furthermore, the methodology developed here is extendable to segmentation tasks for other materials with similar resolution, texture, and contrast characteristics, thereby paving the way for accelerated and automated analysis in corrosion science and beyond.

Artificial Intelligence↗

Estimating energy consumption and GHG emissions in the U.S. food supply chain for net-zero

This work provides a database of the U.S. food system’s energy consumption and GHG emissions at the national and state levels by food supply chain (FSC) stage, fuel type, and food commodity. We estimate that the U.S. FSC consumed a total 4660 TBTU (4900 PJ) of site energy, 7130 TBTU (7500 PJ) of primary energy, and generated 970 MMT of GHG emissions in 2016. Among all the stages, on-farm production is the largest energy consumer (31% primary energy) and GHG emissions contributor (70%), largely due to raising animals. Optimizing distribution can reduce the stage’s energy consumption and GHG emissions and increase products’ shelf-life. Reducing food loss and waste is another good option, as it decreases the amount of food necessary to grow, thus impacting the overall FSC. The database can help stakeholders identify stage- and region-specific strategies and measures to curtail the environmental footprint of the U.S. food system.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Application of machine learning to discover new intermetallic catalysts for the hydrogen evolution and the oxygen reduction reactions

The adsorption energies for hydrogen, oxygen, and hydroxyl were calculated by means of density functional theory on the lowest energy surface of 24 pure metals and 332 binary intermetallic compounds with stoichiometries AB, A 2 B, and A 3 B taking into account the effect of biaxial elastic strains. This information was used to train two random forest regression models, one for the hydrogen adsorption and another for the oxygen and hydroxyl adsorption, based on 9 descriptors that characterized the geometrical and chemical features of the adsorption site as well as the applied strain. All the descriptors for each compound in the models could be obtained from physico-chemical databases. The random forest models were used to predict the adsorption energy for hydrogen, oxygen, and hydroxyl of ≈2700 binary intermetallic compounds with stoichiometries AB, A 2 B, and A 3 B made of metallic elements, excluding those that were environmentally hazardous, radioactive, or toxic. This information was used to search for potential good catalysts for the HER and ORR from the criteria that their adsorption energy for H and O/OH, respectively, should be close to that of Pt. Further, this investigation shows that the suitably trained machine learning models can predict adsorption energies with an accuracy not far away from density functional theory calculations with minimum computational cost from descriptors that are readily available in physico-chemical databases for any compound. Moreover, the strategy presented in this paper can be easily extended to other compounds and catalytic reactions, and is expected to foster the use of ML methods in catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning-enabled discovery of ionic liquid–solvent electrolytes exhibiting high ionic conductivity

Ionic liquids (ILs), which are a class of materials with versatile nature and growing popularity, are facing impediments toward widespread usage as electrolytes due to various factors such as low ionic conductivity, high viscosity, high market price etc. One of the ways these limitations can be addressed is by mixing ILs with a molecular solvent. In a combinatorial sense, there exists an immense number of specific IL–solvent combinations. An exhaustive experimental or even simulation-based investigation of the chemical space spanned by such combinations can be extremely time-consuming, expensive, and nearly impossible. An alternative approach is to employ machine learning-based models developed from available databases. Although there exists prior literature that integrates machine learning to investigate mixtures of specific solvents with ILs, these models lack generalization necessitating development of a large number of ML models to handle various solvents. To remedy this shortcoming, as a part of designing green electrolytes with high ionic conductivity that can have potential applications in next-generation batteries and solar cells, this work aims to develop a unified machine learning model to predict ionic conductivity of any IL–solvent mixture system. In this regard, three models, namely, Random Forest, extreme gradient boosting (XGBoost), and artificial neural network (ANN) were formulated using the NIST ILThermo database. The dataset contained 549 unique ionic liquids from 16 cation families and 81 unique solvents, representing a total of 23 712 datapoints. SHAPLEY additive explanation (SHAP) method was used to assess the impact of various features on model prediction and their significance was compared with literature to gain physical insight about the model behavior. Finally, using the developed models, approximately 2.5 million IL–solvent mixtures at five different compositions were screened at room temperature. The high-throughput screening yielded nearly 19 000 IL–solvent mixtures for which ionic conductivity was found to exceed the ionic conductivity of conventional Li-ion battery electrolyte.

25 ENERGY STORAGE↗

Computational toolkit for predicting thickness of 2D materials using machine learning and autogenerated dataset by large language model

The thickness of 2D materials not only plays a crucial role in determining the performance of nanoelectronic and optoelectronic devices but also introduces complexities in predicting volume-dependent properties, such as energy storage capacity, due to the intrinsic vacuum within these materials. Although a plethora of experimental techniques, including but not limited to optical contrast, Raman spectroscopy, nonlinear optical spectroscopy, near-field optical imaging, and hyperspectral imaging, facilitate the measurement of 2D material thickness, comprehensive data for many materials remain elusive. Over the past decade, the exponential proliferation of 2D materials and their heterostructures has outstripped the capabilities of conventional experimental and computational approaches. In this evolving landscape, machine learning (ML) has emerged as an indispensable tool, offering a scalable approach to augment these traditional methodologies. Addressing the critical gap, we introduce THICK2D—Thickness Hierarchy Inference and Calculation Kit for 2D Materials. This Python-based computational framework harnesses an autogenerated thickness database, developed using large language models, and advanced ML algorithms to facilitate the rapid and scalable estimation of material thickness, relying solely on crystallographic data. To demonstrate the utility and robustness of THICK2D, we successfully used the toolkit to predict the thickness of more than 8000 2D-based materials, sourced from two extensive 2D materials databases. THICK2D is disseminated as an open-source utility, accessible on GitHub at https://github.com/gmp007/THICK2D, and archived on Zenodo at https://10.5281/zenodo.11216648.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Novel application of neutrinos to evaluate U.S. nuclear weapons performance

There is a growing realization that neutrinos can be used as a diagnostic tool to better understand the inner workings of a nuclear weapon. Robust estimates demonstrate that an Inverse Beta Decay (IBD) neutrino scintillation detector built at the Nevada Test Site with a 1000-ton active target mass at a standoff distance of 500 m would detect thousands of antineutrino events per nuclear test. This would provide less than 4% statistical error on the measured antineutrino rate and 5% error on antineutrino energy. Extrapolating this to an error on the test device explosive yield requires knowledge from evaluated nuclear databases, non-equilibrium fission rates, and assumptions on internal neutron fluxes. Initial calculations demonstrate that the total number of neutrinos emitted per fission in the first 10 3 s after a short pulse of 239 Pu fission is about a factor of two less than that from Pu fissioning under steady state conditions. Furthermore, there are significant energy spectral differences as a function of time after the pulse that must be considered. These and other model dependencies will be discussed in the paper. In the absence of nuclear weapons testing, many of the technical and theoretical challenges of a full nuclear test could be mitigated with a low cost smaller scale 20 ton fiducial mass IBD demonstration detector placed near a pulsed reactor. Potential reactors include the Texas A&M University TRIGA 1 GW–10 ms pulsed facility or the Sandia Annular Core Research Reactor. The short duty cycle and repeatability of pulses would provide critical real environment testing and measurements, which would be valuable for planning a possible real test shot in the future. Furthermore, the antineutrino rate as a function of time data would provide unique constraints on fission databases and model assumptions. Finally, there are impactful science drivers such as sensitive searches for ∼1 eV 2 sterile neutrinos and ∼MeV scale axions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗