Search NASASearch

SEARCH · Search NASA

Results for “Databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Database and deep-learning scalability of anharmonic phonon properties by automated brute-force first-principles calculations

Understanding the anharmonic phonon properties of crystal compounds—such as phonon lifetimes and thermal conductivities—is essential for investigating and optimizing their thermal transport behaviors. These properties also impact optical, electronic, and magnetic characteristics through interactions between phonons and other quasiparticles and fields. In this study, we develop an automated first-principles workflow to calculate anharmonic phonon properties and build a comprehensive database encompassing more than 6500 inorganic compounds. Utilizing this dataset, we train a graph neural network model to predict thermal conductivity values and spectra from structural parameters, demonstrating a scaling law in which prediction accuracy improves with increasing training data size. High-throughput screening with the model enables the identification of materials exhibiting extreme thermal conductivities—both high and low. The resulting database offers valuable insights into the anharmonic behavior of phonons, thereby accelerating the design and development of advanced functional materials.

Ohnishi, Masato [University of Tokyo (Japan); Inst

The UCLA Cosmochemistry Database

Abstract The UCLA Cosmochemistry Database was initiated as part of a data-rescue and -storage project aimed at archiving a variety of cosmochemical data acquired at University of California, Los Angeles (UCLA). The data collection includes elemental compositions of extraterrestrial materials analyzed by UCLA cosmochemists over the last five decades. The analytical techniques include atomic absorption spectrometry (AAS) and neutron activation analysis (NAA) at UCLA. The data collection is stored on the Astromaterials Data System (Astromat). We provide both interactive tables and downloadable datasheets for users to access all data. The UCLA Cosmochemistry Database archives cosmochemical data that are essential tools for increasing our understanding of the nature and origin of extraterrestrial materials. Future studies can reference the data collection in the examination, analysis, and classification of newly acquired extraterrestrial samples.

Science & Technology - Other Topics

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),

United States Nuclear Power Reactor Used Nuclear Fuel Database and Applications

The Unified Database (UDB) within STANDARDS serves as the foundational data infrastructure for managing the United States' spent nuclear fuel inventory of 315,111 discharged assemblies totaling 91,036 metric tons of heavy metal. The database organizes this complex inventory through over 200 interconnected tables structured into eight primary attribute categories, supporting integrated analyses across storage, transportation, and disposal domains. Data enters the UDB through the GC-859 Nuclear Fuel Data Survey, which transitioned to web-based collection in 2023, improving data quality through real-time validation. The UDB enables automated generation of input files for nuclear safety analyses, reducing preparation time from weeks to hours while maintaining traceability. Applications include national inventory reporting, Certificate of Compliance assessments, and facility optimization. The three-tier distribution model balances accessibility with security requirements for federal agencies, national laboratories, and research organizations. The UDB provides essential data infrastructure as spent fuel management transitions from site-specific to integrated national campaigns.

Stefanovic, Peter

Assessing the numerical stability of physics models to equilibrium variation through database comparisons on DIII-D

High fidelity kinetic equilibria are crucial for tokamak modeling and analysis. Manual workflows for constructing kinetic equilibria are time consuming and subject to user error, motivating development of automated equilibrium reconstruction tools to provide accurate and consistent reconstructions for downstream physics analysis. These automated tools also provide access to kinetic equilibria at large database scales, which enables the quantification of general uncertainties arising from equilibrium reconstruction techniques. In this paper, we compare a large database of DIII-D kinetic equilibria generated manually by physics experts to equilibria from automated kinetic reconstruction tools, assessing the impact of reconstruction method on equilibrium parameters and resulting magnetohydrodynamic stability calculations. We find agreement among scalar parameters, whereas profile quantities, such as the bootstrap current, show larger disagreements. We analyze ideal kink and classical tearing stability with DCON and STRIDE respectively, finding that the kink stability calculation is generally more robust than the tearing index Δ' calculation. We find that in 90% of cases, both kink stability classifications are unchanged between the manual expert and automated kinetic equilibria.

CAKE

Physics insights from a large-scale 2D UEDGE simulation database for detachment control in KSTAR

A large-scale database of two-dimensional UEDGE simulations has been developed to study detachment physics in KSTAR and to support surrogate models for control applications. Nearly 70 000 steady-state solutions were generated, systematically scanning upstream density, input power, plasma current, impurity fraction, and anomalous transport coefficients, with magnetic and electric drifts across the magnetic field included. The database identifies robust detachment indicators, with strike-point electron temperature at detachment onset consistently T e,target ~ 3-4 eV, largely insensitive to upstream conditions. Scaling relations reveal weaker impurity sensitivity than one-dimensional models and show that heat flux widths follow Eich’s scaling only for uniform, low D and χ. Distinctive in–out divertor asymmetries are observed in KSTAR, differing qualitatively from DIII-D. Complementary time-dependent simulations quantify plasma response to gas puffing, with delays of 5 - 15 ms at the outer strike point and ∼ 40 ms for the low-magnetic-field-side radiation front. These dynamics are well captured by first-order-plus-dead-time models and are consistent with experimentally observed detachment-control behavior in KSTAR (Gupta et al 2025 Plasma Phys. Control. Fusion (submitted)).

Physics - Plasma physics

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES

Genomes OnLine Database (GOLD) v.10: new features and updates

The Genomes OnLine Database (GOLD; https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute is a comprehensive online metadata repository designed to catalog and manage information related to (meta)genomic sequence projects. GOLD provides a centralized platform where researchers can access a wide array of metadata from its four organization levels namely Study, Organism/Biosample, Sequencing Project and Analysis Project. GOLD continues to serve as a valuable resource and has seen significant growth and expansion since its inception in 1997. With its expanded role as a collaborative platform, it not only actively imports data from other primary repositories like National Center for Biotechnology Information but also supports contributions from researchers worldwide. This collaborative approach has enriched the database with diverse datasets, creating a more integrated resource to enhance scientific insights. As genomic research becomes increasingly integral to various scientific disciplines, more researchers and institutions are turning to GOLD for their metadata needs. To meet this growing demand, GOLD has expanded by adding diverse metadata fields, intuitive features, advanced search capabilities and enhanced data visualization tools, making it easier for users to find and interpret relevant information. This manuscript provides an update and highlights the new features introduced over the last 2 years.

59 BASIC BIOLOGICAL SCIENCES

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N

Expansion of the tmRNA sequence database and new tools for search and visualization

Abstract Transfer–messenger RNA (tmRNA) contributes essential tRNA-like and mRNA-like functions during the process of trans-translation, a mechanism of quality control for the translating bacterial ribosome. Proper tmRNA identification benefits the study of trans-translation and also the study of genomic islands, which frequently use the tmRNA gene as an integration site. Automated tmRNA gene identification tools are available, but manual inspection is still important for eliminating false positives. We have increased our database of precisely mapped tmRNA sequences over 50-fold to 97 179 unique sequences. Group I introns had previously been found integrated within a single subsite within the TψC-loop; they have now been identified at four distinct subsites, suggesting multiple founding events of invasion of tmRNA genes by group I introns, all in the same vicinity. tmRNA genes were found in metagenomic archaeal genomes, perhaps a result of misbinning of bacterial sequences during genome assembly. With the expanded database, we have produced new covariance models for improved tmRNA sequence search and new secondary structure visualization tools.

59 BASIC BIOLOGICAL SCIENCES

Resources, Training, and Education Under the Heliostat Consortium: Industry Gap Analysis and Building a Resource Database

Concentrating solar power is not a widely deployed or known technology area, and the heliostat workforce community in the United States is currently small, with knowledge and expertise not widely available. The resource, training, and education (RTE) topic within the Heliostat Consortium (HelioCon) was established to address this. RTE encompasses resources, practices, and programs to ensure that (1) newcomers to the heliostat development community have an adequate knowledge base and training to conduct R&D efforts, (2) outsiders to the field are provided with resources and opportunities to join the workforce, and (3) the workforce community is a productive, healthy, and fulfilling environment for all workers. In the first year of the project, a roadmap study was conducted, in which the major gaps in RTE were identified by consulting experts in the industry, with the top gap being the lack of public accessibility to concentrating solar-thermal power (CSP) knowledge. Here, to address this, the HelioCon team has been developing a centralized web-based resource database, containing a reference library, educational videos, lists of components suppliers and software/metrology tools, a power tower plant database, and information on existing standards/guidelines.

14 SOLAR ENERGY

Utility Rate Database

The Utility Rate Database is a free storehouse of rate structure information from utilities in the United States. The database includes rates for utilities based on the authoritative list of U.S. utility companies maintained by the U.S. Department of Energy’s Energy Information Administration.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

NEWTS State-Level Database Dashboard

The National Energy Water Treatment and Speciation (NEWTS) State-Level Database Dashboard enables quick visualization and exploration of energy-related wastewater data collected from state environmental agencies and research projects across the United States. The dashboard features a subset of the NEWTS Integrated Dataset (version 1.0), which was built to provide stakeholders with a unified and standardized database to support wastewater treatment and critical minerals exploration. The full NEWTS Integrated Dataset and citations for the original data resources can be found here: https://edx.netl.doe.gov/dataset/newts-integrated-dataset-version-1-0

abandoned mine drainage

Advanced Infrastructure Integrity Modeling (AIIM) Onshore Pipeline Database

The Advanced Infrastructure Integrity Modeling (AIIM) Onshore Pipeline Database is an interoperable spatial resource containing critical environmental, operational, and reported stressors tied to publicly available oil and gas pipeline locations across the contiguous U.S. and Alaska. This database contains two layers: 1. Pipeline point locations (‘pipeline_points’) – More than 500,000 points (at every kilometer along pipelines, and end points) to which more than 350 stress-related variables have been appended. 2. Merged pipelines (‘merged_pipelines’) – The original, publicly available pipeline data (see table below) merged together into one feature class.

Carbon Transport

National Energy Water Treatment and Speciation (NEWTS) Database & Dashboard

The Department of Energy's Office of Fossil Energy & Carbon Management (DOE/FECM) through the National Energy Technology Laboratory (NETL) has launched a free online tool, the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard, which can be utilized by community leaders and water researchers to better understand the composition of energy-related wastewater streams. The NEWTS Database and Dashboard provide public access to difficult-to-access datasets, including the original data sources and the processed data forms for input into aqueous chemistry modeling software. The data provided by the tool will help mitigate environmental risks and identify possible sources of valuable critical minerals (CM). The goal of this ASME Power presentation is to highlight the data and capabilities of this free online-tool for obtaining high quality water datasets in formats that are easy for modeling the treatment and recovery of valuable resources from effluent waste stream associated with energy operations.

Siefert, Nicholas

CO2-Locate: A Dynamic Database and Tool for Accessing National Oil and Gas Well Data to Inform Carbon Storage Projects

The CO2-Locate Database is a growing compilation of publicly available wellbore resources that have been merged based on common attributes across data sources with an attribute schema developed to be consistent across disparate resources, reduce data gaps, and eliminate record redundancy. The first version of CO2-Locate has been published to Energy Data eXchange (EDX) and includes the integrated public wells dataset as well as additional geospatial summary layers of key wellbore characteristics to protect proprietary resources. Additionally, the CO2-Locate database has been deployed into a web application, enabling easy access, data filtering capabilities, and visualization of U.S. wellbore infrastructure by stakeholders to inform injection site selection and risk assessments.

Dyer, Alec S. [NETL Site Support Contractor, Natio

National Energy Water Treatment & Speciation (NEWTS): A Water & Critical Mineral Database and Dashboard

The scarcity of water resources, the need for beneficial water reuse, and the challenges of wastewater treatment are becoming increasingly pressing in economic, social, and environmental domains. Addressing these concerns requires effective treatment strategies to manage wastewater streams and tackle environmental and economic issues. Furthermore, the recovery of critical minerals from the waste streams associated with energy production holds the promise of offsetting treatment costs and securing local sources of valuable minerals. However, relevant data on these waste streams are dispersed and challenging to locate. The process of ingesting such data into modeling software often involves multiple steps, requiring data restructuring to meet software-input requirements. The non-standardized reporting of water data makes data aggregation and reformatting a time-consuming process. Additionally, essential attributes necessary for modeling water treatment and mineral scale formation are frequently missing. Moreover, data gaps vary depending on the region of interest. Consequently, there is a pressing need for high-quality energy-water composition data that can be easily imported into water chemistry modeling software. To address this need, the National Energy Technology Laboratory has created the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard—a free online tool catering to community leaders and water researchers. NEWTS facilitates a comprehensive understanding of the composition of energy-related wastewater streams in the United States. The datasets provide detailed concentrations and speciation of major and minor aqueous compounds in energy-related wastewater streams, including power plant leachate, acid mine drainage, brackish water, and oil and gas produced water across the United States. Many of the aqueous species are critical minerals (Li, REEs) in high demand to modernize the world’s energy infrastructure. Many of the datasets also contain volumetric flow-rates needed to model the treatment and reuse scenarios in advanced aqueous chemistry software programs. The NEWTS Database and Dashboard offer public access to hitherto challenging-to-access datasets, presented in a standardized format that is tailored for easy input into aqueous chemistry modeling software. By performing the work needed to transform dispersed, disparate data sources into unified, model-ready datasets, NEWTS serves as an essential resource in advancing water treatment research and sustainable water resource management.

produced water management

Zero Power Reactor Database (ZPRD) Development Plan

Past sodium-cooled fast reactors (SFR) were built with an active experimental program in place to support the design and development work. Most of the experimental facilities in the United States that were important for SFR design were shutdown in the 1980s and 1990s. Reactor licensing and construction requires any reactor design to be verified against existing reactor facilities or experimental measurements. With the absence of those experimental facilities, modern SFR projects must rely on historical measurements to demonstrate that the engineering modeling software and data being used for the new reactor design work are reliable. There has been a considerable push in the last 6 years by both DOE and commercial companies to obtain historical experimental measurements that are relevant for SFRs, in particular those with features that are important for the new reactor designs of interest. The zero power reactor experiments carried out at Argonne National Laboratory’s critical facilities (ZPR-3, ZPR-6, ZPR-9, and ZPPR) from the 1950s to the 1980s are some of the best reactor physics experiments on SFR technology that are available today. Of particular interest today are the ZPPR-15 measurements done at the ZPPR facility for the Integral Fast Reactor project in the 1980s as they are in line with most commercial and DOE interests today. In the past 10 years, the measurements done on ZPPR-15 have been processed into both Monte Carlo (MCNP) and deterministic models (MC2-3 and DIF3D) useable for validating the engineering modeling software for key parts of the SFR design work. To achieve this, a detailed model description must be created for the experiment and the experimental measurement that the engineering modeling software is to reproduce. Then, an assessment of the uncertainty on the measured quantity which considers all of the sources of uncertainty in defining the model must be obtained and documented. The models created for ZPPR-15 provide the best validation basis available today for neutronics modeling software. Reference 2 is a good resource to understand how these models were built and how the uncertainties on the measured quantities were derived. The intention of the Zero Power Reactor Database (ZPRD), hosted at frdb.ne.anl.gov, is to make available the experimental measurements and models that have been constructed to-date. Though ZPPR-15 measurements are the primary data requested for validation needs, other measurements on ZPPR, ZPR-6, and ZPR-9 in support of the Clinch River Breeder Reactor (CRBR) and Fast Test Reactor (FFTF) should also be considered important for future software validation needs. In this manuscript, the details of available measurements on ZPR-3, ZPR-6, ZPR-9, and ZPPR facilities are summarized, and a general organization of the web interface is displayed. Many of the documents associated with the measurements are export controlled information so access to the database will also have to be controlled.

22 GENERAL STUDIES OF NUCLEAR REACTORS