Search NASA⌕ Search

SEARCH · Search NASA

Results for “DATABASE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic in-context learning with conversational models for data extraction and materials property prediction

The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challenge. In this paper, we introduce PropertyExtractor, an open-source tool that leverages advanced conversational LLMs such as Google gemini-pro and OpenAI gpt-4, blends zero-shot with few-shot in-context learning, and employs engineered prompts for the dynamic refinement of structured information hierarchies—enabling autonomous, efficient, scalable, and accurate identification, extraction, and verification of material property data. Our tests on material data demonstrate precision and recall that exceed 95% with an error rate of ∼9%, highlighting the effectiveness and versatility of the toolkit. Finally, databases for 2D material thicknesses, a critical parameter for device integration, and energy bandgap values are developed using PropertyExtractor. In particular, for the thickness database, the rapid evolution of the field has outpaced both experimental measurements and computational methods, creating a significant data gap. Our work addresses this gap and showcases the potential of PropertyExtractor as a reliable and efficient tool for the autonomous generation of various material property databases, advancing the field.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Analysis of the impact of parallel magnetic fluctuations on linear gyrokinetic stability in NSTX-U and verification of gyro-fluid models

In this work, we use the CGYRO gyrokinetic code to analyze two L- and one H-mode discharges from the National Spherical Torus Experiment (NSTX) and NSTX-Upgrade (NSTX-U) selected due to their different mix of ion-scale driftwaves, ion temperature gradient (ITG) mode and trapped electron mode (TEM), and electromagnetic instabilities, kinetic ballooning mode (KBM), and micro-tearing mode (MTM) in the plasma core. It is found that the effect of parallel magnetic fluctuations is strongly destabilizing to the unstable KBMs compared to calculations with only perpendicular magnetic fluctuations. Two discharges have a mix of ITG/TEM and MTMs that are predicted to be dominant instability across the plasma radius. The parallel magnetic fluctuations are found to have little effect on the MTM stability but are destabilizing to ITG/TEM modes. To test the validity of the gyro-fluid linear stability codes TGLF and GFS at low aspect ratio, a database of linear growth rates has been created using the CGYRO gyrokinetic code. The database is comprised of various parameter scans around a standardized set of NSTX-U core parameters. It contains a group of electrostatic cases and an electromagnetic group that includes the effects of perpendicular and parallel magnetic fluctuations. Comparing the results from the GFS and TGLF models, we find that GFS exhibits the best agreement with the database of CGYRO linear growth rates. Comparing the model results for the electromagnetic scans shows that GFS captures the effects of parallel magnetic fluctuations accurately, while the TGLF model does not, as it lacks sufficient perpendicular energy resolution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Prediction of carbon nanostructure mechanical properties and the role of defects using machine learning

Graphene-based nanostructures hold immense potential as strong and lightweight materials, however, their mechanical properties such as modulus and strength are difficult to fully exploit due to challenges in atomic-scale engineering. This study presents a database of over 2,000 pristine and defective nanoscale CNT bundles and other graphitic assemblies, inspired by microscopy, with associated stress–strain curves from reactive molecular dynamics (MD) simulations using the reactive INTERFACE force field (IFF-R). These 3D structures, containing up to 80,000 atoms, enable detailed analyses of structure-stiffness-failure relationships. By leveraging the database and physics- and chemistry-informed machine learning (ML), accurate predictions of elastic moduli and tensile strength are demonstrated at speeds 1,000 to 10,000 times faster than efficient MD simulations. Hierarchical Graph Neural Networks with Spatial Information (HS-GNNs) are introduced, which integrate chemistry knowledge. HS-GNNs as well as extreme gradient boosted trees (XGBoost) achieve forecasts of mechanical properties of arbitrary carbon nanostructures with only 3 to 6% mean relative error. The reliability equals experimental accuracy and is up to 20 times higher than other ML methods. Predictions maintain 8 to 18% accuracy for large CNT bundles, CNT junctions, and carbon fiber cross-sections outside the training distribution. The physics- and chemistry-informed HS-GNN works remarkably well for data outside the training range while XGBoost works well with limited training data inside the training range. The carbon nanostructure database is designed for integration with multimodal experimental and simulation data, scalable beyond 100 nm size, and extendable to chemically similar compounds and broader property ranges. The ML approaches have potential for applications in structural materials, nanoelectronics, and carbon-based catalysts.

Winetrout, Jordan J.↗

Validation of Fast Reactor Depletion Tools Using EBR-II Measured Data

The validation of simulation tools for calculating fuel depletion and evolution in fast reactors is vital for design, licensing, deployment, operations, and material accountancy. The Physics Analysis Database (PADB) and Analytical Laboratory (AL) database contain measured data collected from Experimental Breeder Reactor II (EBR-II) and were used to validate the most recent versions of the Argonne Reactor Computation (ARC) tool suite and ORIGEN-S for calculating isotopic compositions in irradiated fast reactor fuel. The PADB contains important modeling and operational information about the EBR-II core design, fuel cycle, and analytical results from the legacy versions of the ARC tool suite. The AL database contains the measured isotopic compositions of irradiated samples taken from core subassemblies. A new procedure was developed for the ARC tool suite to perform the EBR-II depletion simulation, as well as to perform more detailed isotopic calculations using ORIGEN-S calculations by coupling it with the ARC suite. Both the ARC and ARC-ORIGEN results were compared with the AL measured data for all relevant samples and showed good agreement for the major actinides. Good agreement with measured data was also achieved using the ARC-ORIGEN approach for several fission products.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Metrics and extrapolation of resonant magnetic perturbation thresholds for ELM suppression

This large database study of resonant magnetic perturbation (RMP) edge localized mode (ELM) suppression thresholds in the AUG, DIII-D, EAST, and KSTAR tokamaks details the key strengths and weaknesses of RMP metrics. The RMP ELM suppression database used for this work contains plasma information at the time of transition from ELMing to ELM suppressed states where a clear experimental threshold is identified. The experimental threshold distributions are compared for five metrics: (1) the island overlap width, (2) pedestal top Chirikov overlap, (3) peeling edge displacement, (4) pedestal top resonant drive, and (5) edge dominant mode overlap. The distributions, the regularity of the dependence on RMP coil currents, and the sensitivities of a given metric to equilibrium reconstruction details are compared. The overlap metric proves to be a good compromise between including the appropriate plasma response physics and maintaining a numerical robustness. This quantity does not exhibit clear power-law scalings for projection, but machine learning can assist in predicting thresholds within the existing parameter ranges and providing uncertainty quantification of those predictions. Two new first-principles models, one utilizing a threshold from the non-linear Modified Rutherford equation evaluated at the pedestal top and one utilizing the SLAYER code to calculate the linear tearing threshold from torque balance, offer possible paths to extrapolation beyond the existing database parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL↗

Uncertainties in greenhouse gas emission factors: A comprehensive analysis of switchgrass‐based biofuel production

Abstract This study investigates uncertainties in greenhouse gas (GHG) emission factors related to switchgrass‐based biofuel production in Michigan. Using three life cycle assessment (LCA) databases—US lifecycle inventory (USLCI) database, GREET, and Ecoinvent—each with multiple versions, we recalculated the global warming intensity (GWI) and GHG mitigation potential in a static calculation. Employing Monte Carlo simulations along with local and global sensitivity analyses, we assess uncertainties and pinpoint key parameters influencing GWI. The convergence of results across our previous study, static calculations, and Monte Carlo simulations enhances the credibility of estimated GWI values. Static calculations, validated by Monte Carlo simulations, offer reasonable central tendencies, providing a robust foundation for policy considerations. However, the wider range observed in Monte Carlo simulations underscores the importance of potential variations and uncertainties in real‐world applications. Sensitivity analyses identify biofuel yield, GHG emissions of electricity, and soil organic carbon (SOC) change as pivotal parameters influencing GWI. Decreasing uncertainties in GWI may be achieved by making greater efforts to acquire more precise data on these parameters. Our study emphasizes the significance of considering diverse GHG factors and databases in GWI assessments and stresses the need for accurate electricity fuel mixes, crucial information for refining GWI assessments and informing strategies for sustainable biofuel production.

Kim, Seungdo↗

Catalog of topological phonon materials

Phonons play a crucial role in many properties of solid-state systems, and it is expected that topological phonons may lead to rich and unconventional physics. On the basis of the existing phonon materials databases, we have compiled a catalog of topological phonon bands for more than 10,000 three-dimensional crystalline materials. Using topological quantum chemistry, we calculated the band representations, compatibility relations, and band topologies of each isolated set of phonon bands for the materials in the phonon databases. Additionally, we calculated the real-space invariants for all the topologically trivial bands and classified them as atomic or obstructed atomic bands. We have selected more than 1000 “ideal” nontrivial phonon materials to motivate future experiments. The datasets were used to build the Topological Phonon Database.

Science & Technology - Other Topics↗

Collection And Analysis Of Telemetry For The Cyote Heuristic

CATCH CLI focuses on gathering telemetry data, storing it in the Neo4j database, querying for Mitre ATT&CK patterns, and creating STIX 2.1 reports. Key Components: Analysis Modules: Analyze data to detect attack patterns. GoSTOTS Collection Engines: Collect telemetry data. These tools can be used together or individually. Analysis modules rely on data from specific engines to identify attack patterns. Source Code Organization: Engines: CATCH/catch/cmd/collection Modules: CATCH/catch/cmd/analysis CGUI Overview CATCH Graphical User Interface (CGUI) offers a graphical shell to execute CATCH CLI, allowing easy editing of: Analysis Modules Database configurations Profiles (collection and device settings) Neo4j Overview Neo4j is a graph database using the Cypher query language, storing data in JSON. It seamlessly integrates with STIX 2.1 data for: Data Submission: CATCH Collection Engines Data Querying: Analysis Modules CATCH modifies STIX 2.1 data for Neo4j submission and reverts it back during querying. STIG Overview Structured Threat Intelligence Graph (STIG) is a tool for creating, editing, querying, analyzing, and visualizing threat intelligence using STIX 2.1 and storing data in Neo4j. Usage Tools can be run: Manually (CLI): Refer to CATCH documentation User Interface: Run ./cgui/CGUI or go run ./cgui/ Additional Information Logging System: Detailed in the config documentation Further Documentation: Available for CATCH and CGUI

Madsen, MichaelJ. [Idaho National Laboratory (INL)↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

Responsive Assistant For Navigating And Guiding Engineering With Rigor (ranger)

The bot uses the GitHub API to fetch discussions from a MOOSE repository and store the data in a vector database. When a new discussion is initiated, the algorithm compares the discussion title with the content of all previous discussions (title + discussions) in the database and provides the most relevant posts to the user. The database is updated regularly to include all new posts, potentially on a monthly basis.

Li, Mengnan [Idaho National Laboratory (INL), Idah↗

datasight [SWR-26-045]

This software is an AI-powered data exploration with natural language. datasight connects an AI agent to your database and provides a web UI where you can ask questions in natural language. The agent writes SQL, runs queries, and generates interactive Plotly visualizations. Supports DuckDB, PostgreSQL, SQLite, and Flight SQL databases. Also queries local CSV and Parquet files directly — no database setup required. Supports Anthropic Claude (default), GitHub Models (open source), and Ollama (local) as LLM backends.

Thom, Daniel [National Laboratory of the Rockies (↗

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence↗

Untargeted, tandem mass spectrometry (LC/MS-MS) metaproteomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization. This package contains soil metaproteomics data in the context of site specific metagenomes from soil depth profiles in three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. These metaproteomes were collected in 2018 after 4.5 years of warming from five depth intervals (0-10 cm, 10-30 cm, 30-45 cm, 45-60 cm, 60-80 cm). For protein identification, the collected spectra were searched following a target-decoy search strategy against a database of metagenome predicted proteins (covering 96 samples from 2014 to 2021) representing the complete sequence diversity at the site. Data was searched with mass spectrometry database search tool (MS-GF+) using Pacific Northwest National Laboratory (PNNL)'s Data Management System (DMS) Processing pipeline. The metagenomes are published as part of another data package. Raw metaproteomic data and the data products from MS-GF+ are deposited in the Mass Spectrometry Interactive Virtual Environment (MassIVE) database under accession no. MSV000097826. Here we present a dataset that includes spectral counts for the detected proteins across samples (EMSL50964_BrodieAllMAGs_Globals_SC.txt), the sequences of the detected proteins, and sample metadata file that contains site information for the soil metaproteome samples.

Belowground Biogeochemistry Science Focus Area↗

NEWTS Integrated Dataset (version 1.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v1.0 provides water researchers, community leaders, and regulators with a unified and standardized energy-related wastewater stream database. This resource is derived from 27 state and federal entities, and scientific publications, and contains more than 400,000 sample records, many of which also provide geospatial information. The dataset includes data for several different energy-related wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, and power plant wastewater. The NEWTS Integrated Dataset was built to support environmentally prudent decision-making, explore treatment opportunities, and identify potential critical mineral sources. A subset of this novel resource is also featured on NETL NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

abandoned mine drainage↗

NEWTS Integrated Dataset (version 2.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v2.0 provides water researchers, community leaders, regulators, and industry stakeholders with a unified and standardized energy-process wastewater chemistry database. This resource is derived from 39 state and federal entities, and scientific publications, and contains more than 700,000 sample records, many of which also provide geospatial information. The dataset includes chemistry data for several different energy-process wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, power plant wastewater, and geothermal fluids. The NEWTS Integrated Dataset was built to support prudent decision-making, characterization of potential critical mineral sources, and modeling of treatment and valorization options. A subset of this novel resource is also featured on the NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

AMD↗

Chemical Recommender System: Replacement Suggestions for Small Molecules

The Chemical Recommender System (CRS) is an open-source, high-performance toolkit that enables real-time similarity searches across the complete PubChem database (over 50 million molecules) using commodity hardware. The CRS addresses critical limitations in existing chemical informatics platforms through a novel vector database infrastructure, extensible model integration capabilities, and complete algorithmic transparency. The system implements a vector database deployment with partitioned indexing that achieves a ~60x speedup over traditional approaches. A containerized model integration framework allows researchers to seamlessly incorporate custom predictive models into the full-scale search and scoring pipeline, while complete configurability of search parameters, filtering logic, and scoring functions provides capabilities not available in existing black-box solutions. Beyond structural similarity, the CRS integrates OPERA QSAR models for thermophysical and toxicity predictions, RDKit synthetic accessibility scoring, and user-defined models to compute weighted final replacement scores. The complete system is accessible through an interactive web application supporting real-time progress monitoring, post-processing score re-weighting, automated PDF reporting, and batch processing capabilities.

Nair, Parthiv Anand [Sandia National Laboratories ↗