Search NASA⌕ Search

SEARCH · Search NASA

Results for “Databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

A Proxy Method to Bridge LCA Data Gaps Using Automated Material Classification and Probabilistic Under-Specification

Life cycle assessments (LCAs) are essential for understanding the environmental impacts of material production. However, gaps in life cycle inventory (LCI) data for material and chemical inputs present a key challenge for LCA practitioners, especially in the early design stages. Strategies for filling in these gaps require additional time and expertise, which can hinder the LCA’s completion. This study combined automatic material classification and probabilistic under-specification to create a time-efficient method to fill material LCI data gaps. To illustrate the proposed method, proxy environmental impact distributions were generated using publicly available material LCI data classified into the ChemOnt chemical taxonomy using the open-source chemical classification software ClassyFire. Input materials with data gaps were then classified into the same taxonomy, where proxy environmental impact values could be selected from the available distributions to quickly fill in any data gaps. Although these methods were applied to classify material production processes available in the Federal LCA Commons and Ecoinvent databases, they can be applied to any LCA database. This study shows that classifying materials by their chemical structure produces taxonomies with increased granularity relative to industrial classification, improving the ability of under-specified proxy data to be used for differentiating the environmental impacts of competing designs.

biological databases↗

Twins in rotational spectroscopy: Does a rotational spectrum uniquely identify a molecule?

Rotational spectroscopy is the most accurate method for determining structures of molecules in the gas phase. It is often assumed that a rotational spectrum is a unique “fingerprint” of a molecule. The availability of large molecular databases and the development of artificial intelligence methods for spectroscopy make the testing of this assumption timely. In this paper, we pose the determination of molecular structures from rotational spectra as an inverse problem. Within this framework, we adopt a funnel-based approach to search for molecular twins, which are two or more molecules, which have similar rotational spectra but distinctly different molecular structures. Here we demonstrate that there are twins within standard levels of computational accuracy by generating rotational constants for many molecules from several large molecular databases, indicating that the inverse problem is ill-posed. However, some twins can be distinguished by increasing the accuracy of the theoretical methods or by performing additional experiments.

74 ATOMIC AND MOLECULAR PHYSICS↗

Simulation Tools for Characterizing Stress Distribution in Laser Welded Dissimilar Joints

This project focuses on developing a thermo-metallurgical-mechanical modeling method to accurately predict the microstructural evolution and residual stress in laser welding between dissimilar metals, such as HSLA steel and high carbon equivalent (CE) gear steel. The method leverages a comprehensive material database to model the temperature and rate dependent phase transformations, along with their associated effects on material properties, such as thermal expansion and flow stress, throughout the welding process. A key innovation is the incorporation of phase transformation and phase-specific properties, which enhances the accuracy of residual stress predictions. The mixture material in the fusion zone due to the dissimilar metals will also be addressed in the numerical model. This is especially critical in scenarios involving phase transformations in the fusion zone and heat-affected zone (HAZ), where the phase changes can induce substantial residual stress variations. The material database has been generated using JMatPro. The modeling approach is implemented through a custom User Material (UMAT) subroutine, executed with the commercial finite element software Abaqus.

36 MATERIALS SCIENCE↗

A Play-Based Exploration of CO 2 Storage in the Illinois Basin (Final Technical Report)

This report documents the results of "A Play-Based Exploration of CO₂ Storage in the Illinois Basin" (DE-FE0032366), a project funded by the U.S. Department of Energy Office of Fossil Energy and Carbon Management and conducted by the Illinois State Geological Survey (ISGS) at the University of Illinois Urbana-Champaign in partnership with Visage Energy. The project adapted play-based exploration (PBE), a systematic basin-scale evaluation methodology from the petroleum industry, to screen areas of Illinois for commercial geologic carbon storage (GCS) in Cambro-Ordovician strata. The traditional play concept was expanded to encompass three play element groups, subsurface geologic factors, surface features, and societal factors, yielding 24 play elements with defensible suitability criteria applied through a five-tier classification scheme. An integrated geospatial database was assembled from ISGS, MGSC, MRCI, NATCARB, and public data sources, supported by significant data-improvement work including correction of legacy well locations, digitization of more than 1,300 well construction records using the DOE CATALOG team's OGRRE tool, compilation of a statewide 2D seismic database, and production of a refined fault and fold geodatabase.

58 GEOSCIENCES↗

Mining Thermophile Photosynthesis Genes: A Synthetic Operon Expressing Chloroflexota Species Reaction Center Genes in Rhodobacter sphaeroides

Photosynthesis is the foundation of the vast majority of life systems, and is therefore the most important bioenergetic process on earth. The greatest diversity of photosynthetic systems is found in microorganisms. However, our understanding of the biophysical and biochemical processes that transduce light into chemical energy is derived from a relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of predicted proteins that catalyze the initial photochemical reactions deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for the expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42–90 °C) springs available through the JGI IMG database, to generate a resource of diverse sequences that are potentially adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota↗

SEAFORML (Smart Exploration and Analysis For Optimal and Robust Machine Learning)

The poster discusses data analysis of the WAVgraph database and applied machine learning methods for it. The database is a long-term project that seeks to be a comprehensive repository of information on cyber threats and is updated regularly. It was previously unanalyzed and unexplored. The goal was to learn more about it and its contents in order to have a better understanding and enable better use. The data analysis and discovery enabled further exploration through natural language processing, similarity, and clustering methods. The poster shows some of the insights from the analysis and explains the methods used for the machine learning applications.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

MaTableGPT: GPT‐Based Table Data Extractor from Materials Science Literature

Abstract Efficiently extracting data from tables in the scientific literature is pivotal for building large‐scale databases. However, the tables reported in materials science papers exist in highly diverse forms; thus, rule‐based extractions are an ineffective approach. To overcome this challenge, the study presents MaTableGPT, which is a GPT‐based table data extractor from the materials science literature. MaTableGPT features key strategies of table data representation and table splitting for better GPT comprehension and filtering hallucinated information through follow‐up questions. When applied to a vast volume of water splitting catalysis literature, MaTableGPT achieves an extraction accuracy (total F1 score) of up to 96.8%. Through comprehensive evaluations of the GPT usage cost, labeling cost, and extraction accuracy for the learning methods of zero‐shot, few‐shot, and fine‐tuning, the study presents a Pareto‐front mapping where the few‐shot learning method is found to be the most balanced solution owing to both its high extraction accuracy (total F1 score >95%) and low cost (GPT usage cost of 5.97 US dollars and labeling cost of 10 I/O paired examples). The statistical analyses conducted on the database generated by MaTableGPT revealed valuable insights into the distribution of the overpotential and elemental utilization across the reported catalysts in the water splitting literature.

Yi, Gyeong Hoon [Computational Science Research Ce↗

Fragme∩t: An Open‐Source Framework for Multiscale Quantum Chemistry Based on Fragmentation

Fragment-based quantum chemistry offers a means to circumvent the nonlinear computational scaling of conventional electronic structure calculations, by partitioning a large calculation into smaller subsystems then considering the many-body interactions between them. Variants of this approach have been used to parameterize classical force fields and machine learning potentials, applications that benefit from interoperability between quantum chemistry codes. However, there is a dearth of software that provides interoperability yet is purpose-built to handle the combinatorial complexity of fragment-based calculations. To fill this void we introduce “Fragme∩t”, an open-source software application that provides a tool for community validation of fragment-based methods, a platform for developing new approximations, and a framework for analyzing many-body interactions. Fragme∩t includes algorithms for automatic fragment generation and structure modification, and for distance- and energy-based screening of the requisite subsystems. Checkpointing, database management, and parallelization are handled internally and results are archived in a portable database. Interfaces to various quantum chemistry engines are easy to write and exist already for Q-Chem, PySCF, xTB, Orca, CP2K, MRCC, Psi4, NWChem, GAMESS, and MOPAC. Applications reported here demonstrate parallel efficiencies around 96% on more than 1000 processors but also showcase that the code can handle large-scale protein fragmentation using only workstation hardware, all with a codebase that is designed to be usable by non-experts. Fragme∩t conforms to modern software engineering best practices and is built upon well established technologies including Python, SQLite, and Ray. The source code is available under the Apache 2.0 license.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interaction between the emerging components of online shopping and in-person activities: insights from a behavioral survey

The rise of technological advancements has led to the commonplace practice of online shopping for retail, grocery, and food. However, little research has been conducted on the interplay of these components in burdened communities (BCs) that face issues of marginalization and limited access to digital resources. Here, this study aims to provide a comprehensive understanding of travel behavior changes by analyzing the interconnectedness of the emerging components of online shopping (retail, grocery, and food) and in-person activities in both BCs and non-BCs. A unique household-level database is created by linking the 2021 Puget Sound Household Travel Survey and the US Department of Transportation’s burdened community databases, and a conditional mixed process model is estimated to account for unobserved endogeneity. The findings suggest households living in BCs are less likely to order online retail goods and groceries compared to non-BC households. Additionally, the probability of making more restaurant trips decreases for households living in BCs. The study highlights the digital divide that exists in BCs and the differences in online and in-person shopping activities across socioeconomic levels. Policymakers may address these disparities to promote better access to goods and services for all. Besides, planners may need to improve the travel demand models by accounting for the emerging components of online shopping and the trip frequencies by purpose in BCs.

Digital Divide↗

Implementation of an extensible property modeling framework in ESPEI with applications to molar volume and elastic stiffness models

Property models are becoming more widely adopted by commercial Calphad databases, but they are not nearly as common in non-commercial or traditional academic Calphad databases. A primary driver is that user-friendly Calphad modeling tools that support property models are not widely available. Here we present new property modeling capabilities that have been implemented in ESPEI (the Extensible, Self-optimizing Phase Equilibrium Infrastructure). These capabilities include both generating property model parameters from data and improvements to the algorithmic selection of the most appropriate model from a series of candidates. Additionally, two illustrative examples are given that use ESPEI to fit different property models. First, we generate molar volume model parameters for Group IV, V, and VI refractory BCC alloys based on the model by Lu et al. (2005). Second, we demonstrate the extensibility of ESPEI’s property modeling capabilities by implementing a custom PyCalphad model for BCC elastic stiffness parameters to generate and compare parameters to the ones assessed by Marker et al. (2018) using the same data. Property models generated by ESPEI can be used in PyCalphad or further optimized with uncertainty quantification using ESPEI.

36 MATERIALS SCIENCE↗

Dataset describing two reference models for full-spectral lighting and daylight simulations together with implementations for two software systems

A dataset of two spectral lighting simulation reference models - one office and one factory hall - is presented. It aims to demonstrate and support full-spectral daylight and electric lighting simulations and facilitate evaluation of non-visual effects of light. The dataset includes Rhino CAD geometry, comprehensive spectral material and light source data and window system BSDF data. Example implementations in the two software tools, Radiance and OWL, enable reproducible workflows and support adoption in other software. The dataset is openly available on Zenodo. The office model reproduces Room 518 at the University of Innsbruck, including a west-facing façade and interior furnishings. The factory hall model follows the proposed geometry in the European standard 15193 for building energy performance. Interior reflectances in the office were measured in-situ using a handheld spectrometer. Exterior spectra and factory hall materials matching specified reflectances were obtained from an online spectral materials database. Glazing transmittance was derived from IGDB data using LBNL Optics/WINDOW. BSDFs for venetian blinds at various tilt angles, and for a diffusing pane adapted from the Complex Glazing Database, were generated in WINDOW. Luminaires in both models are specified with photometric files (Eulumdat/IES) and lamp spectra (Fluorescent 840, 4000 K LED). The provided example implementations (Radiance, OWL) include prepared input data and scripts to run first spectral simulations; example results are also included. The dataset is prepared to support reuse by researchers, designers and software developers for method validation, software engineering and comparison, and development of spectral metrics and controls.

Geisler-Moroder, David↗

A generative machine learning model for designing metal hydrides applied to hydrogen storage

Developing new metal hydrides is a critical step toward efficient hydrogen storage in carbon-neutral energy systems. However, existing materials databases, such as the Materials Project, contain a limited number of well-characterized hydrides, which constrains the discovery of optimal candidates. This work presents a framework that integrates causal discovery with a lightweight generative machine learning model to generate novel metal hydride candidates that may not exist in current databases. Using a dataset of 450 samples (270 training, 90 validation, and 90 testing), the model generates 1000 candidates. After ranking and filtering, six previously unreported chemical formulas and crystal structures are identified, four of which are validated by density functional theory simulations and show strong potential for future experimental investigation. Overall, the proposed framework provides a scalable and time-efficient approach for expanding hydrogen storage datasets and accelerating materials discovery.

generative model↗

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES↗

Mobility assessment of the BCC and carbide phases in the C-Nb, C-U and Nb-U systems

Uranium carbides with refractory metal additions are considered for Gen IV nuclear reactors and nuclear thermal propulsion as fuels for their high-temperature and corrosion resistant properties. Understanding kinetic effects that dictate microstructural evolution during fabrication and operating conditions is essential to advance technological development of these fuels. This work presents the development of an atomic mobility database for C-Nb-U systems based off available experimental data supported with ab-initio methods. The mobility assessments and uncertainty quantification (using Markov chain Monte Carlo) were conducted in the Kawin software. Carbon diffusion is considered dominant, as metal diffusion is much slower, with niobium diffusion being even slower and rate limiting than uranium metal. We provide a comprehensive and self-consistent thermo-kinetic database that is validated by diffusion couple simulations through Kawin. In conclusion, this enables prediction of microstructural and phase evolution critical for the development and lifetime assessment of next generation nuclear fuels.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Identification of drug repurposing candidates for amyotrophic lateral sclerosis using electronic health records: a retrospective cohort study

Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disease with a life expectancy of only 3–5 years and few approved treatments. To identify drug repurposing candidates for the treatment of ALS, we analysed the electronic health records (EHRs) of a large cohort of military veterans with ALS. We analysed the EHRs of individuals in the US Veterans Health Administration (VHA) database who were diagnosed with ALS between Jan 1, 2009 and Dec 31, 2019 to assess medication effects. Individuals without recorded prescriptions after the date of diagnosis were excluded. Two sets of criteria were applied to ascertain exposure. Exposure criteria A were met if the dispense date or the end date of the medication was within 12 months of ALS diagnosis and the end date was at least 6 months after the dispense date. Exposure criteria B were met if there were at least two dispenses within 6 months before diagnosis and 12 months after diagnosis. Propensity score-matched control groups were generated on the basis of confounders included in the EHR, with methodology of potential outcomes used to infer treatment effects. The primary outcome was death. A standard Cox proportional hazards analysis was done to assess association with survival. Survival was defined as the time from diagnosis date recorded in the EHR to death reported in the Department for Veterans Affairs Vital Status File. Follow-up survival time was censored on Dec 31, 2020, for those alive on this date. Downstream protein targets of drugs with clinically significant effects were analysed using the protein–protein interaction networks-based algorithm PathFX. The EHRs of 11 003 individuals with ALS in the VHA database were appropriate for analysis. 162 medications with treatment groups of 30 or more individuals were identified. Among these 162 medications, 27 were associated with statistically significant changes (≥0·1) in the hazard ratio (HR) for death. 18 of the medications were associated with a reduced HR for death (prolonged survival), and nine were associated with an increased HR for death (reduced survival). Drugs associated with reduced HR included HMG-CoA reductase inhibitors (simvastatin, pravastatin, lovastatin, and atorvastatin), PDE5 inhibitors (vardenafil and sildenafil), and α-adrenergic antagonists (tamsulosin and terazosin). The medications associated with an increased HR were drugs used either in the management of clinical features of ALS associated with poor outcomes or in end-of-life care. PathFx analysis identified a complex of proteins interacting with several of the identified drugs. To our knowledge, this analysis is the largest EHR-based study for identifying drug repurposing candidates for ALS. We identified several drugs that warrant further assessment as therapeutic options in ALS, as well as a protein network complex that might serve as a therapeutic target for ALS.

Reimer, Richard J. [Stanford Univ., CA (United Sta↗

Particle control via cryopumping and its impact on the edge plasma profiles of Alcator C-Mod

At the high n e proposed for high-field fusion reactors, it is uncertain whether ionization, as opposed to plasma transport, will be most influential in determining n e at the pedestal and separatrix. A database of Alcator C-Mod discharges is analyzed to evaluate the impact of source modification via cryopumping. The database contains similarly-shaped H-modes at fixed I P = 0.8 MA and B t = 5.4 T, spanning a large range in P net and ionization. Measurements from an edge Thomson scattering system are combined with those from a midplane-viewing Ly α camera to evaluate changes to n e and T e in response to changes in ionization rates, S ion ∙ $n^{sep}_e$ and $T^{ped}_e$ are found to be most sensitive to changes to $S^{sep}_{ion}$, as opposed to $n^{ped}_e$ and $T^{sep}_e$. Dimensionless quantities, namely α MHD and v*, are found to regulate attainable pedestal values. Select discharges at different values of P net and in different pumping configurations are analyzed further using SOLPS-ITER. It is determined that changes to plasma transport coefficients are required to self-consistently model both plasma and neutral edge dynamics. Pumping is found to modify the poloidal distribution of atomic neutral density, n 0 , along the separatrix, increasing n 0 at the active X-point. Opaqueness to neutrals from high n e in the divertor is found to play a role in mediating neutral penetration lengths and hence, the poloidal distribution of neutrals along the separatrix. Pumped discharges thus require a larger particle diffusion coefficient than that inferred purely from 1D experimental profiles at the outer midplane.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Development of a BISON validation case for the TRISO transient irradiations in NSRR using effective heat capacity methods

The current tristructural isotropic (TRISO) fuel assessment and validation database in BISON primarily covers steady- state irradiation and high-temperature furnace testing. Transient assessment cases are potentially needed to support U.S. industry efforts in designing and deploying commercial reactors using TRISO fuels. Historical transient tests in- volving TRISO fuels used highly conservative conditions compared to the typical high-temperature gas-cooled reactor accident scenarios. Despite this, modeling historical transient tests is fundamental for evaluating BISON’s predictive capabilities, adapting material properties for high-temperature and high-particle-power regimes, and developing a sys- tematic validation approach for TRISO transient applications. This work developed a 1D model of transient experiments carried out at the Nuclear Safety Research Reactor using BISON. BISON’s predictions of energy deposition, UO 2 melting onset, and molten volume fractions are compared against experimental measurements. Melting was modeled using an effective specific heat capacity model for UO 2 . We found that BISON’s predictions are in reasonable agreement with experimental data for low-energy-deposition cases, and that BISON overpredicts melting at higher energy depositions. We also discuss the potential causes of discrepancies between the simulated and measured results and propose ways to further develop the model. Although these simulations used conservative conditions compared to those expected for actual TRISO-fueled reactors, they extended the range of conditions reflected in the data in the existing BISON database for TRISO fuels.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Annotation of DOM metabolomes with an ultrahigh resolution mass spectrometry molecular formula library

Current approaches to analyzing metabolomic data often rely on matching MS/MS fragmentation data to sparse libraries or databases. This approach results in limited identification of features, often with less than 10% of the dataset being annotated. A complementary approach is to assign molecular formula to features based on accurate mass measurements, but the platforms commonly used for metabolomics do not have the needed accuracy or resolving power to do this robustly, particularly for larger molecules. Using our newly modified analysis tool, CoreMS, we generated a library of molecular formula from pooled samples analyzed with LC-21T FT-ICR MS. This library successfully annotated approximately 53.2% of features identified from the exometabolome of marine diatom Phaeodactylum tricornutum – a nearly ten-fold increase over the 5.9% annotation rate achieved using a conventional MS/MS library matching approach. Using this FT-ICR MS library approach, we were able to differentiate differences in the exometabolome of P. tricornutum in iron replete and iron limited conditions, with 668 metabolites being differentially expressed (p < 0.05, 2 x intensity difference) under these conditions. The traditional MS/MS fragmentation-based annotation approach only annotated 61 of these metabolites, while our novel pipeline annotated 450 metabolites and revealed 12 metabolites that were significantly more abundant under low iron conditions. Our results demonstrate the utility of ultrahigh resolution mass spectrometry for generating more comprehensive and confident molecular annotations.

21T-FTICR-MS, CoreMS↗