Search NASA⌕ Search

SEARCH · Search NASA

Results for “DATABASE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Database-Agnostic Log Analysis and Monitoring Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, US, IL; Fermilab]↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

Carbon Management Projects (CONNECT) Database and Explorer

Overview The Carbon Management Projects (CONNECT) Toolkit is an online exploratory visualization tool developed by the U.S. Department of Energy's (DOE) Office of Fossil Energy and Carbon Management (FECM) with support from other federal agencies such as the U.S. Environmental Protection Agency (EPA) and the U.S. Department of Transportation (DOT). It provides a single point of access to authoritative information on federal agency investment in a portfolio of research, development, and demonstration (RD&D) projects that have been publicly announced to advance technologies for point source carbon capture, carbon dioxide removal, transport, storage, and conversion, collectively referred to as carbon management. The RD&D programs covered in this tool are authorized by annual congressional appropriations ("Base Program") and the 2021 Infrastructure Investment and Jobs Act (IIJA). The tool also incorporates public information on other federal initiatives, such as the Regional Clean Hydrogen Hubs, and public information released by other government agencies, such as the Environmental Protection Agency's (EPA) and Primacy States’ Underground Injection Control Class VI permits and EPA’s facility level greenhouse gas (GHG) emissions. Developed in a geographic information system, the tool organizes carbon management projects into five groups based on the primary technology that a project aims to advance, each visually represented as a digital layer ("carbon management project layer"). Only federally funded projects are included, which can be awarded projects that are completed or ongoing, or projects that have been selected but are currently under negotiation. Project information can be viewed in the map or in the attribute table below it when turned on. In the map view, each project is displayed at either its host site (for field work), where available, or its performer site (project lead's location, further explained in the table below). Host sites and performer sites are represented in distinct icons. Several reference layers offer additional public information on infrastructural and natural resource environment for carbon management. These reference layers, combined with multiple geographical basemaps, enable users to visualize the carbon management project layers in context. Carbon management project information will be updated monthly based on feedback and information availability. Carbon management project layers Point Source Carbon Capture (PSC) This layer contains DOE-funded projects focused on capturing carbon dioxide (CO2) from power plants or industrial facilities. Carbon Dioxide Removal (CDR) This layer contains DOE-funded projects focused on capturing CO2 from the atmosphere, including direct air capture (DAC) and DAC hubs, direct ocean capture, enhanced mineralization, and biomass carbon removal and storage. For projects with multiple host sites, each of the sites are displayed individually with the project cost and cost sharing information representing the total for the entire project. Carbon Transport This layer contains DOE- and DOT-funded projects focused on CO2 transport. The Transport Research and Development sublayer contains projects that do not involve physical infrastructure; the Proposed Transport Corridor sublayer contains projects for which either a route for the transport infrastructure has been proposed or a general area for the transport infrastructure has been identified. Carbon Storage This layer contains DOE-funded key projects focused on CO2 storage. For projects with multiple field-work sites, each of the sites are displayed individually on the map with the project cost and cost sharing information representing the overall total for the entire project. Carbon Conversion This layer contains DOE-funded projects focused on converting CO2 into economically valuable products. Reference layers The following layers provide additional information in the geographic proximity of carbon management projects. Users should reference the original sources for more details (weblinks provided below and in pop-up windows on the map). Regional Clean Hydrogen Hub and Facility These layers illustrate the approximate areas of the Regional Clean Hydrogen Hubs announced by DOE's Office of Clean Energy Demonstrations (OCED) and the approximate locations of individual facilities that constitute the hubs (see "Where are the H2Hubs located?" on the webpage linked above). EPA Facility Level GHG Emissions (direct emitter) This layer shows direct CO2 emissions from stationary sources in 2022, using data extracted from EPA's Facility Level Information on GreenHouse gases Tool (FLIGHT). Captured and injected CO2 are not deducted from direct emitters’ total emissions. Contact EPA for additional details. Underground Injection Control Class VI permit/permit application This layer shows the locations of CO2 injection wells that are granted or in the process of applying for an Underground Injection Control Class VI permit by EPA or a Primacy State (currently Louisiana, North Dakota, and Wyoming). The URLs for the permits or permit applications are provided in the pop-up windows associated with the well locations. Contact EPA for additional details. Carbon Storage Resource This layer contains information on prospective CO2 storage resources in saline formations and oil and gas reservoirs provided by the National Carbon Sequestration Database and Geographic Information System (NATCARB) spatial database. Contact NETL for additional details. Existing CO2 pipeline This layer shows active CO2 pipelines based on information digitized from the map issued by the Pipeline and Hazardous Materials Safety Administration (PHMSA). Contact PHMSA for additional details.

Carbon Conversion↗

Database and deep-learning scalability of anharmonic phonon properties by automated brute-force first-principles calculations

Understanding the anharmonic phonon properties of crystal compounds—such as phonon lifetimes and thermal conductivities—is essential for investigating and optimizing their thermal transport behaviors. These properties also impact optical, electronic, and magnetic characteristics through interactions between phonons and other quasiparticles and fields. In this study, we develop an automated first-principles workflow to calculate anharmonic phonon properties and build a comprehensive database encompassing more than 6500 inorganic compounds. Utilizing this dataset, we train a graph neural network model to predict thermal conductivity values and spectra from structural parameters, demonstrating a scaling law in which prediction accuracy improves with increasing training data size. High-throughput screening with the model enables the identification of materials exhibiting extreme thermal conductivities—both high and low. The resulting database offers valuable insights into the anharmonic behavior of phonons, thereby accelerating the design and development of advanced functional materials.

Ohnishi, Masato [University of Tokyo (Japan); Inst↗

The UCLA Cosmochemistry Database

Abstract The UCLA Cosmochemistry Database was initiated as part of a data-rescue and -storage project aimed at archiving a variety of cosmochemical data acquired at University of California, Los Angeles (UCLA). The data collection includes elemental compositions of extraterrestrial materials analyzed by UCLA cosmochemists over the last five decades. The analytical techniques include atomic absorption spectrometry (AAS) and neutron activation analysis (NAA) at UCLA. The data collection is stored on the Astromaterials Data System (Astromat). We provide both interactive tables and downloadable datasheets for users to access all data. The UCLA Cosmochemistry Database archives cosmochemical data that are essential tools for increasing our understanding of the nature and origin of extraterrestrial materials. Future studies can reference the data collection in the examination, analysis, and classification of newly acquired extraterrestrial samples.

Science & Technology - Other Topics↗

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),↗

United States Nuclear Power Reactor Used Nuclear Fuel Database and Applications

The Unified Database (UDB) within STANDARDS serves as the foundational data infrastructure for managing the United States' spent nuclear fuel inventory of 315,111 discharged assemblies totaling 91,036 metric tons of heavy metal. The database organizes this complex inventory through over 200 interconnected tables structured into eight primary attribute categories, supporting integrated analyses across storage, transportation, and disposal domains. Data enters the UDB through the GC-859 Nuclear Fuel Data Survey, which transitioned to web-based collection in 2023, improving data quality through real-time validation. The UDB enables automated generation of input files for nuclear safety analyses, reducing preparation time from weeks to hours while maintaining traceability. Applications include national inventory reporting, Certificate of Compliance assessments, and facility optimization. The three-tier distribution model balances accessibility with security requirements for federal agencies, national laboratories, and research organizations. The UDB provides essential data infrastructure as spent fuel management transitions from site-specific to integrated national campaigns.

Stefanovic, Peter↗

Assessing the numerical stability of physics models to equilibrium variation through database comparisons on DIII-D

High fidelity kinetic equilibria are crucial for tokamak modeling and analysis. Manual workflows for constructing kinetic equilibria are time consuming and subject to user error, motivating development of automated equilibrium reconstruction tools to provide accurate and consistent reconstructions for downstream physics analysis. These automated tools also provide access to kinetic equilibria at large database scales, which enables the quantification of general uncertainties arising from equilibrium reconstruction techniques. In this paper, we compare a large database of DIII-D kinetic equilibria generated manually by physics experts to equilibria from automated kinetic reconstruction tools, assessing the impact of reconstruction method on equilibrium parameters and resulting magnetohydrodynamic stability calculations. We find agreement among scalar parameters, whereas profile quantities, such as the bootstrap current, show larger disagreements. We analyze ideal kink and classical tearing stability with DCON and STRIDE respectively, finding that the kink stability calculation is generally more robust than the tearing index Δ' calculation. We find that in 90% of cases, both kink stability classifications are unchanged between the manual expert and automated kinetic equilibria.

CAKE↗

Physics insights from a large-scale 2D UEDGE simulation database for detachment control in KSTAR

A large-scale database of two-dimensional UEDGE simulations has been developed to study detachment physics in KSTAR and to support surrogate models for control applications. Nearly 70 000 steady-state solutions were generated, systematically scanning upstream density, input power, plasma current, impurity fraction, and anomalous transport coefficients, with magnetic and electric drifts across the magnetic field included. The database identifies robust detachment indicators, with strike-point electron temperature at detachment onset consistently T e,target ~ 3-4 eV, largely insensitive to upstream conditions. Scaling relations reveal weaker impurity sensitivity than one-dimensional models and show that heat flux widths follow Eich’s scaling only for uniform, low D and χ. Distinctive in–out divertor asymmetries are observed in KSTAR, differing qualitatively from DIII-D. Complementary time-dependent simulations quantify plasma response to gas puffing, with delays of 5 - 15 ms at the outer strike point and ∼ 40 ms for the low-magnetic-field-side radiation front. These dynamics are well captured by first-order-plus-dead-time models and are consistent with experimentally observed detachment-control behavior in KSTAR (Gupta et al 2025 Plasma Phys. Control. Fusion (submitted)).

Physics - Plasma physics↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes OnLine Database (GOLD) v.10: new features and updates

The Genomes OnLine Database (GOLD; https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute is a comprehensive online metadata repository designed to catalog and manage information related to (meta)genomic sequence projects. GOLD provides a centralized platform where researchers can access a wide array of metadata from its four organization levels namely Study, Organism/Biosample, Sequencing Project and Analysis Project. GOLD continues to serve as a valuable resource and has seen significant growth and expansion since its inception in 1997. With its expanded role as a collaborative platform, it not only actively imports data from other primary repositories like National Center for Biotechnology Information but also supports contributions from researchers worldwide. This collaborative approach has enriched the database with diverse datasets, creating a more integrated resource to enhance scientific insights. As genomic research becomes increasingly integral to various scientific disciplines, more researchers and institutions are turning to GOLD for their metadata needs. To meet this growing demand, GOLD has expanded by adding diverse metadata fields, intuitive features, advanced search capabilities and enhanced data visualization tools, making it easier for users to find and interpret relevant information. This manuscript provides an update and highlights the new features introduced over the last 2 years.

59 BASIC BIOLOGICAL SCIENCES↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

Expansion of the tmRNA sequence database and new tools for search and visualization

Abstract Transfer–messenger RNA (tmRNA) contributes essential tRNA-like and mRNA-like functions during the process of trans-translation, a mechanism of quality control for the translating bacterial ribosome. Proper tmRNA identification benefits the study of trans-translation and also the study of genomic islands, which frequently use the tmRNA gene as an integration site. Automated tmRNA gene identification tools are available, but manual inspection is still important for eliminating false positives. We have increased our database of precisely mapped tmRNA sequences over 50-fold to 97 179 unique sequences. Group I introns had previously been found integrated within a single subsite within the TψC-loop; they have now been identified at four distinct subsites, suggesting multiple founding events of invasion of tmRNA genes by group I introns, all in the same vicinity. tmRNA genes were found in metagenomic archaeal genomes, perhaps a result of misbinning of bacterial sequences during genome assembly. With the expanded database, we have produced new covariance models for improved tmRNA sequence search and new secondary structure visualization tools.

59 BASIC BIOLOGICAL SCIENCES↗

Resources, Training, and Education Under the Heliostat Consortium: Industry Gap Analysis and Building a Resource Database

Concentrating solar power is not a widely deployed or known technology area, and the heliostat workforce community in the United States is currently small, with knowledge and expertise not widely available. The resource, training, and education (RTE) topic within the Heliostat Consortium (HelioCon) was established to address this. RTE encompasses resources, practices, and programs to ensure that (1) newcomers to the heliostat development community have an adequate knowledge base and training to conduct R&D efforts, (2) outsiders to the field are provided with resources and opportunities to join the workforce, and (3) the workforce community is a productive, healthy, and fulfilling environment for all workers. In the first year of the project, a roadmap study was conducted, in which the major gaps in RTE were identified by consulting experts in the industry, with the top gap being the lack of public accessibility to concentrating solar-thermal power (CSP) knowledge. Here, to address this, the HelioCon team has been developing a centralized web-based resource database, containing a reference library, educational videos, lists of components suppliers and software/metrology tools, a power tower plant database, and information on existing standards/guidelines.

14 SOLAR ENERGY↗

Utility Rate Database

The Utility Rate Database is a free storehouse of rate structure information from utilities in the United States. The database includes rates for utilities based on the authoritative list of U.S. utility companies maintained by the U.S. Department of Energy’s Energy Information Administration.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

NEWTS State-Level Database Dashboard

The National Energy Water Treatment and Speciation (NEWTS) State-Level Database Dashboard enables quick visualization and exploration of energy-related wastewater data collected from state environmental agencies and research projects across the United States. The dashboard features a subset of the NEWTS Integrated Dataset (version 1.0), which was built to provide stakeholders with a unified and standardized database to support wastewater treatment and critical minerals exploration. The full NEWTS Integrated Dataset and citations for the original data resources can be found here: https://edx.netl.doe.gov/dataset/newts-integrated-dataset-version-1-0

abandoned mine drainage↗

Advanced Infrastructure Integrity Modeling (AIIM) Onshore Pipeline Database

The Advanced Infrastructure Integrity Modeling (AIIM) Onshore Pipeline Database is an interoperable spatial resource containing critical environmental, operational, and reported stressors tied to publicly available oil and gas pipeline locations across the contiguous U.S. and Alaska. This database contains two layers: 1. Pipeline point locations (‘pipeline_points’) – More than 500,000 points (at every kilometer along pipelines, and end points) to which more than 350 stress-related variables have been appended. 2. Merged pipelines (‘merged_pipelines’) – The original, publicly available pipeline data (see table below) merged together into one feature class.

Carbon Transport↗

Developing a Database of Bio-based Materials for Building Envelope Applications

Oak Ridge National Laboratory (ORNL) has been funded by the Department of Energy (DOE) to help accelerate the introduction of building envelope materials that would reduce the carbon footprint of the buildings sector. The DOE’s Building Technologies Office has historically sought to resolve the knowledge gaps regarding the energy efficiency and moisture durability of building envelope systems and to develop the data, guidance, and tools needed to facilitate rapid industry adoption of high-performance, moisture-managed envelope systems. This project will help accelerate the widespread acceptance of a new generation of building materials developed specifically with the intent of reducing the carbon footprint of buildings. We have produced a database of hygrothermal transport properties on low embodied carbon building materials that can be added to energy and durability simulation tools. Properties that were measured include density, heat capacity, thermal conductivity as a function of temperature and relative humidity, moisture dependent permeance, and sorption isotherms as a function of relative humidity. These data sets were measured following consensus national standards using state-of-the-art facilities. The data has been compiled and is being made available to building designers who require these data to assess these new materials in their designs. We will publish the data and seek its addition to reference databases such as the ASHRAE Handbook of Fundamentals.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗