Search NASA⌕ Search

SEARCH · Search NASA

Results for “Databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Uncertainties in greenhouse gas emission factors: A comprehensive analysis of switchgrass‐based biofuel production

Abstract This study investigates uncertainties in greenhouse gas (GHG) emission factors related to switchgrass‐based biofuel production in Michigan. Using three life cycle assessment (LCA) databases—US lifecycle inventory (USLCI) database, GREET, and Ecoinvent—each with multiple versions, we recalculated the global warming intensity (GWI) and GHG mitigation potential in a static calculation. Employing Monte Carlo simulations along with local and global sensitivity analyses, we assess uncertainties and pinpoint key parameters influencing GWI. The convergence of results across our previous study, static calculations, and Monte Carlo simulations enhances the credibility of estimated GWI values. Static calculations, validated by Monte Carlo simulations, offer reasonable central tendencies, providing a robust foundation for policy considerations. However, the wider range observed in Monte Carlo simulations underscores the importance of potential variations and uncertainties in real‐world applications. Sensitivity analyses identify biofuel yield, GHG emissions of electricity, and soil organic carbon (SOC) change as pivotal parameters influencing GWI. Decreasing uncertainties in GWI may be achieved by making greater efforts to acquire more precise data on these parameters. Our study emphasizes the significance of considering diverse GHG factors and databases in GWI assessments and stresses the need for accurate electricity fuel mixes, crucial information for refining GWI assessments and informing strategies for sustainable biofuel production.

Kim, Seungdo↗

Catalog of topological phonon materials

Phonons play a crucial role in many properties of solid-state systems, and it is expected that topological phonons may lead to rich and unconventional physics. On the basis of the existing phonon materials databases, we have compiled a catalog of topological phonon bands for more than 10,000 three-dimensional crystalline materials. Using topological quantum chemistry, we calculated the band representations, compatibility relations, and band topologies of each isolated set of phonon bands for the materials in the phonon databases. Additionally, we calculated the real-space invariants for all the topologically trivial bands and classified them as atomic or obstructed atomic bands. We have selected more than 1000 “ideal” nontrivial phonon materials to motivate future experiments. The datasets were used to build the Topological Phonon Database.

Science & Technology - Other Topics↗

Collection And Analysis Of Telemetry For The Cyote Heuristic

CATCH CLI focuses on gathering telemetry data, storing it in the Neo4j database, querying for Mitre ATT&CK patterns, and creating STIX 2.1 reports. Key Components: Analysis Modules: Analyze data to detect attack patterns. GoSTOTS Collection Engines: Collect telemetry data. These tools can be used together or individually. Analysis modules rely on data from specific engines to identify attack patterns. Source Code Organization: Engines: CATCH/catch/cmd/collection Modules: CATCH/catch/cmd/analysis CGUI Overview CATCH Graphical User Interface (CGUI) offers a graphical shell to execute CATCH CLI, allowing easy editing of: Analysis Modules Database configurations Profiles (collection and device settings) Neo4j Overview Neo4j is a graph database using the Cypher query language, storing data in JSON. It seamlessly integrates with STIX 2.1 data for: Data Submission: CATCH Collection Engines Data Querying: Analysis Modules CATCH modifies STIX 2.1 data for Neo4j submission and reverts it back during querying. STIG Overview Structured Threat Intelligence Graph (STIG) is a tool for creating, editing, querying, analyzing, and visualizing threat intelligence using STIX 2.1 and storing data in Neo4j. Usage Tools can be run: Manually (CLI): Refer to CATCH documentation User Interface: Run ./cgui/CGUI or go run ./cgui/ Additional Information Logging System: Detailed in the config documentation Further Documentation: Available for CATCH and CGUI

Madsen, MichaelJ. [Idaho National Laboratory (INL)↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

Responsive Assistant For Navigating And Guiding Engineering With Rigor (ranger)

The bot uses the GitHub API to fetch discussions from a MOOSE repository and store the data in a vector database. When a new discussion is initiated, the algorithm compares the discussion title with the content of all previous discussions (title + discussions) in the database and provides the most relevant posts to the user. The database is updated regularly to include all new posts, potentially on a monthly basis.

Li, Mengnan [Idaho National Laboratory (INL), Idah↗

datasight [SWR-26-045]

This software is an AI-powered data exploration with natural language. datasight connects an AI agent to your database and provides a web UI where you can ask questions in natural language. The agent writes SQL, runs queries, and generates interactive Plotly visualizations. Supports DuckDB, PostgreSQL, SQLite, and Flight SQL databases. Also queries local CSV and Parquet files directly — no database setup required. Supports Anthropic Claude (default), GitHub Models (open source), and Ollama (local) as LLM backends.

Thom, Daniel [National Laboratory of the Rockies (↗

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence↗

Untargeted, tandem mass spectrometry (LC/MS-MS) metaproteomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization. This package contains soil metaproteomics data in the context of site specific metagenomes from soil depth profiles in three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. These metaproteomes were collected in 2018 after 4.5 years of warming from five depth intervals (0-10 cm, 10-30 cm, 30-45 cm, 45-60 cm, 60-80 cm). For protein identification, the collected spectra were searched following a target-decoy search strategy against a database of metagenome predicted proteins (covering 96 samples from 2014 to 2021) representing the complete sequence diversity at the site. Data was searched with mass spectrometry database search tool (MS-GF+) using Pacific Northwest National Laboratory (PNNL)'s Data Management System (DMS) Processing pipeline. The metagenomes are published as part of another data package. Raw metaproteomic data and the data products from MS-GF+ are deposited in the Mass Spectrometry Interactive Virtual Environment (MassIVE) database under accession no. MSV000097826. Here we present a dataset that includes spectral counts for the detected proteins across samples (EMSL50964_BrodieAllMAGs_Globals_SC.txt), the sequences of the detected proteins, and sample metadata file that contains site information for the soil metaproteome samples.

Belowground Biogeochemistry Science Focus Area↗

NEWTS Integrated Dataset (version 1.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v1.0 provides water researchers, community leaders, and regulators with a unified and standardized energy-related wastewater stream database. This resource is derived from 27 state and federal entities, and scientific publications, and contains more than 400,000 sample records, many of which also provide geospatial information. The dataset includes data for several different energy-related wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, and power plant wastewater. The NEWTS Integrated Dataset was built to support environmentally prudent decision-making, explore treatment opportunities, and identify potential critical mineral sources. A subset of this novel resource is also featured on NETL NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

abandoned mine drainage↗

NEWTS Integrated Dataset (version 2.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v2.0 provides water researchers, community leaders, regulators, and industry stakeholders with a unified and standardized energy-process wastewater chemistry database. This resource is derived from 39 state and federal entities, and scientific publications, and contains more than 700,000 sample records, many of which also provide geospatial information. The dataset includes chemistry data for several different energy-process wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, power plant wastewater, and geothermal fluids. The NEWTS Integrated Dataset was built to support prudent decision-making, characterization of potential critical mineral sources, and modeling of treatment and valorization options. A subset of this novel resource is also featured on the NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

AMD↗

Chemical Recommender System: Replacement Suggestions for Small Molecules

The Chemical Recommender System (CRS) is an open-source, high-performance toolkit that enables real-time similarity searches across the complete PubChem database (over 50 million molecules) using commodity hardware. The CRS addresses critical limitations in existing chemical informatics platforms through a novel vector database infrastructure, extensible model integration capabilities, and complete algorithmic transparency. The system implements a vector database deployment with partitioned indexing that achieves a ~60x speedup over traditional approaches. A containerized model integration framework allows researchers to seamlessly incorporate custom predictive models into the full-scale search and scoring pipeline, while complete configurability of search parameters, filtering logic, and scoring functions provides capabilities not available in existing black-box solutions. Beyond structural similarity, the CRS integrates OPERA QSAR models for thermophysical and toxicity predictions, RDKit synthetic accessibility scoring, and user-defined models to compute weighted final replacement scores. The complete system is accessible through an interactive web application supporting real-time progress monitoring, post-processing score re-weighting, automated PDF reporting, and batch processing capabilities.

Nair, Parthiv Anand [Sandia National Laboratories ↗

Sierra/SD - Its2Sierra - User's Manual - 5.20

The Integrated Tiger Series (ITS) generates a database containing energy deposition data. This data, when stored on an Exodus file, is not typically suitable for analysis within SierraMechanics for finite element analysis. The its2sierra tool maps data from the ITS database to the Sierra database. This document provides information on the usage of its2sierra.

97 MATHEMATICS AND COMPUTING↗

Data Visualization and Analytics for Optimal Process Parameter Selection in Turning

The objective of this project is to research physics-guided machine learning methods to recommend optimal tools and machining process parameters for turning applications using the MSC test database. For a given turning application, the MSC metalworking specialist needs to make decisions on tools and the associated process parameters for the MSC customer. For a given material, there are many alternatives for tools and a wide range of process parameters to consider. MSC has built a database of tools and parameters for different applications from the historical turning tests completed at various customer sites. The research project aims to use machine learning methods to predict optimal tools and process parameters for the MSC metalworking specialists using the MSC test database. This enables continuous learning of optimal tool and process parameters for different applications as new information is collected from testing. Through MSC, this information can be shared with machining shops across the US leading to improved productivity and efficiency.

42 ENGINEERING↗

M3SF-24LL010301062-Surface Complexation/Ion Exchange Data Integration for Radionuclide Sorption to Clay Minerals

This progress report (Level 3 Milestone Number M3SF-24LL010301062) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Argillite International Collaborations Work Package SF-24LL01030106. The activity is focused on our long-term commitment to engaging our partners in international nuclear waste repository research. The focus of this milestone is the establishment of international collaborations for sorption modeling and the associated impacts of unlocking larger, community-based datasets. More specifically, we are developing a database framework for Spent Fuel and Waste and Science Technology (SFWST) that is aligned with the Helmholtz Zentrum Dresden Rossendorf (HZDR) and other international sorption database development groups in support of the database needs of the SFWST program.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Results from an Aeromagnetic Survey to Detect Steel-Cased Wells at a Marcellus Shale Well Site in Washington County, Pennsylvania

Pennsylvania has a 150-year history of oil and gas production—the longest of any state—and this enduring activity has resulted in the drilling of more than 300,000 recorded wells. However, unknown wells likely exist because innumerable wells were drilled during Pennsylvania’s intense early oil and gas history when incomplete records were kept of well locations. There is concern that early wells are likely to be ineffectively sealed because there were no laws that required plugging when the wells were abandoned. Today, many undocumented and unplugged wells are thought to be in areas of emerging shale gas and shale oil development where open wellbores can provide a pathway for undesired upward migration of fluids and gas from hydraulically fractured reservoirs. Due to this concern, Pennsylvania regulators have asked operators to locate orphaned and abandoned wells within a 1,000-ft buffer of proposed new wells. The objective of this report is to demonstrate that high-resolution aeromagnetic surveys, historic air photos, and Light Detection and Ranging (LiDAR) imagery can be rapid and effective methods to reconnoiter large, forested areas of moderate terrain for the presence of abandoned wells. These well-finding methods were evaluated at a proposed Marcellus Shale gas drilling site in Washington County, Pennsylvania, where the methods collectively located 18 confirmed wells: 15 wells were identified from aeromagnetic surveys, two wells were identified from inspection of historical air photos, and one well was identified by evaluation of state-wide LiDAR imagery. Only six wells were previously known, and their locations, as recorded in Pennsylvania’s statewide oil and gas wells database (PA/IRIS/WIS), were often too inaccurate for the wells to be found in the dense underbrush. Twelve wells identified in this study were abandoned, unmarked, and undocumented. Aeromagnetic surveys locate wells by detecting the unique magnetic signature of vertical, steel well casing, which is depicted on magnetic maps as a “bull’s eye” type anomaly that is centered directly over the well. However, when wells were drilled and found to be sub-economic, their casing was sometimes pulled and salvaged for reuse. Such wellbores provide no magnetic response and go undetected if all casing was removed. Oftentimes attempts to retrieve well casing were not 100% successful. For example, historical records for one well in the study area indicate that the well was completed in 1902 as a dry hole and that, to the extent possible, the casing was pulled for reuse. However, a section of 10-in. diameter steel casing was not recovered and remains at an unknown depth in the wellbore. This well was easily detected by the aeromagnetic survey although only deep casing remained in the well. To mitigate for the likelihood that wellbores exist where most or all casing has been removed, this study augmented aeromagnetic data with historic air photos and digital terrain models generated from LiDAR datasets—both databases are publicly available at no cost for areas within Pennsylvania. These complementary methods located three wells where the aeromagnetic anomaly, although present, was subtle and overlooked. Together, these methods determined accurate locations for six known wells within the study area and located 12 previously unknown wells. Although it is not certain that these methods successfully located all wells in the study area, the application of these methods does represent a significant improvement over relying on existing databases for well locations. For the Appendix to the report, see: https://www.netl.doe.gov/energy-analysis/details?id=b46c417a-7c9e-4d25-b810-e6248b0217f4</p>

04 OIL SHALES AND TAR SANDS↗

Sierra/SD – Its2Sierra – User's Manual – 5.22

The Integrated Tiger Series (ITS) generates a database containing energy deposition data. This data, when stored on an Exodus file, is not typically suitable for analysis within Sierra Mechanics for finite element analysis. The its2sierra tool maps data from the ITS database to the Sierra database. This document provides information on the usage of its2sierra.

97 MATHEMATICS AND COMPUTING↗

Sierra/SD – Its2Sierra – User’s Manual (V.5.24)

The Integrated Tiger Series (ITS) generates a database containing energy deposition data. This data, when stored on an Exodus file, is not typically suitable for analysis within Sierra Mechanics for finite element analysis. The its2sierra tool maps data from the ITS database to the Sierra database. This document provides information on the usage of its2sierra.

97 MATHEMATICS AND COMPUTING↗

TIGER, A thermodynamics Equilibrium Tool for Explosives Update: Improved Solver

The thermodynamic equilibrium code known as TIGER has been in use since the 1970s and is designed to calculate the performance of energetic materials during detonation at high temperatures and pressures. The original TIGER code utilized a limited database consisting of 12 gaseous and 3 condensed constituents, which were composed of the elements carbon (C), hydrogen (H), nitrogen (N), oxygen (O), and aluminum (Al). In contrast, the more modern JCZS3 database features a significantly larger dataset, including 756 gaseous species, 189 positive and negative ions, and 496 condensed constituents derived from 62 different elements. While this expanded product species database allows for the exploration of a broader range of problems relevant to our laboratory, it also introduces challenges related to convergence, particularly when both liquid and solid species coexist under a vapor dome. In this work, we present an improved TIGER solution technique aimed at accurately determining the equilibrium state for systems that can produce a diverse array of products, including both solid and liquid condensed species. This report presents the status of the TIGER solver as of the end of FY25. Ongoing efforts are focused on enhancing the solver, and the report outlines several planned improvements. We anticipate that further developments will require additional documentation in the future.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗