Search NASASearch

SEARCH · Search NASA

Results for “relational databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

SoK: What does it Mean to Benchmark Database Forensics?

Relational Database Management Systems are the backbone of modern enterprises and public-sector services, and are thus frequent targets of security incidents, insider threats, and thorough regulatory audits. Consequently, databases have become key sources of digital evidence, requiring investigators to reconstruct past activity from audit logs, transaction logs, and backups. Although benchmarking frameworks such as those developed by the Transaction Processing Performance Council (TPC) are widely used to evaluate database performance, they do not capture forensic requirements such as evidentiary completeness, tamper-evidence, chain of custody, or regulatory compliance under GDPR and CCPA. This survey examines the emerging domain of forensic database benchmarking. We gathered prior research on database forensics, secure logging, and tamper-evident data structures; we analyze modern forensic-ready features in commercial and open-source systems (SQL Server Ledger, Oracle Blockchain Tables, PostgreSQL pgAudit, Db2 Audit, Aurora Database Activity Streams, Oracle Real Application Security and IBM Guardium) and assess why existing benchmarks are insufficient. We propose forensic workloads, metrics, and methodologies that incorporate adversarial stressors, deleted-record recovery, and backup analysis. We also identify open research problems and call for a community-driven forensic benchmark suite. The result is an idea for evaluating not only database performance but also forensic soundness, bridging the gap between system engineering, compliance, and digital investigations.

Lenard, Ben

Specifications of Legacy U(Pu)Zr Metallography Data

The DOE Nuclear Energy Advanced Reactor Technologies (ART) Program has supported the creation of several databases with information describing the safety performance of fast reactors, components, and fuels. This growing collection of legacy experimental data, operating data, and analysis is available online to registered users. Metallography data is one of the most important types of PIE data being collected, organized and stored in several ART Fast Reactor Databases (https://frdb.ne.anl.gov), including the Fuels Irradiation & Physics Database (FIPD), Out-of-Pile Transient Database (OPTD), and TREAT (the Transient Reactor Test Facility) Experimental Relational Database (TREXR). These databases contain three main sets of metallography data. The first set is the metallography data from Experimental Breeder Reactor-II (EBR-II) and Fast Flux Test Facility (FFTF) irradiated fuel pins measured in the Hot Fuel Examination Facility (HFEF); the second set is the metallography data from EBR-II irradiated fuel pins measured in the Alpha-Gamma Hot Cell Facility (AGHCF); the third set is the metallography data from transient tested fuel pins (the transients tests include the out-of-pile tests and TREAT tests) measured in AGHCF. Since both the second and third sets of data were measured in AGHCF, they are governed by the same specification. The metallography data in the databases are digital images scanned from either positive or negative photos. The quality of the images, including their resolution and contrast, relies on the preserved quality of the pictures and the scanning conditions. To analyze the microstructure of a fuel pin, a series of preparatory steps must be undertaken. These include sectioning, epoxy mounting, mechanical grinding and polishing, and often etching. The resulting samples were then transferred to a secondary hot cell, or glovebox (depending the strength of radiation field) for microscopy examination. The specifications provided herein focus on the metallography examinations. The sample grinding, polishing and etching processes are also discussed. The procedures involving sectioning and epoxy mounting are out of the scope of the current specification; details can be found in the corresponding operation manuals. If more data are collected and added to the ART Fast Reactor Databases, this specification will be updated to accommodate the additional data.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

HIV Molecular Immunology 2025

HIV Molecular Immunology is a companion volume to HIV Sequence Compendium. This publication, the 2025 edition, is the PDF version of Los Alamos Na tional Laboratory’s web-based HIV Molecular Immunology Database (https://www.hiv.lanl.gov/content/ immunology/). The web interface for this relational database has many search interfaces for HIV immunological in formation, as well as interactive tools to help immunologists design reagents and interpret their results.

59 BASIC BIOLOGICAL SCIENCES

Determination of Scale Bar for AGHCF Metallography Data

The DOE Nuclear Energy Advanced Reactor Technologies (ART) Fast Reactor Program (FRP) has supported the development of several databases containing information on the safety performance of fast reactors, components, and fuels. This growing collection of legacy experimental data, operating data, and analysis is available online to registered users. Metallography data represents one of the most critical types of post-irradiation examination (PIE) data being collected, organized, and archived in several ART Fast Reactor Databases (https://frdb.ne.anl.gov), including the Metallic Fuels Irradiation & Physics Database (FIPD), Out-of-Pile Transient Database (OPTD), and TREAT (the Transient Reactor Test Facility) Experimental Relational Database (TREXR). These databases contain three principal sets of metallography data. The first set comprises metallography data from Experimental Breeder Reactor-II (EBR-II) and Fast Flux Test Facility (FFTF) irradiated fuel pins examined in the Hot Fuel Examination Facility (HFEF). The second set consists of metallography data from EBR-II irradiated fuel pins examined in the Alpha-Gamma Hot Cell Facility (AGHCF). The third set includes metallography data from transient-tested fuel pins (including both out-of-pile furnace tests and TREAT tests) examined in AGHCF. Since both the second and third sets were generated in AGHCF, they are governed by identical specifications. The metallography data in the databases consist of digital images scanned from either positive or negative photographic films. To analyze the microstructure of a fuel pin, a series of preparatory steps are required, including sectioning, epoxy mounting, mechanical grinding and polishing, and etching. Following sample preparation, specimens are transferred for metallographic examination. The AGHCF and HFEF metallography data were generated using optical microscopes manufactured by Leitz and Bausch and Lomb (B&L). Images were recorded on Polaroid film at preset magnifications. Magnification verification for the Leitz and B&L metallographs was conducted every two months prior to 1989 and at least every six months from 1989 through the conclusion of the IFR program. Magnifications determined from imaging of microslide standards were compared to the instrument settings for magnifications ranging from 50× to 500×. If the magnifications determined from standards deviated from the instrument settings, adjustments were made to the bellows extension until agreement was achieved. The specifications for AGHCF and HFEF legacy metallography data have been established based on available hard-copy and digital records, most of which have been incorporated into the data repositories associated with FIPD, OPTD, and TREXR. Detailed specifications including hard-copy records, digitized records, cutting diagrams and sectioning schemes, high-magnification photographs, photomosaics (composites), information tags, scale bars, and magnification verification procedures can be found in a separate report.

22 GENERAL STUDIES OF NUCLEAR REACTORS

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics

Cataloging Legacy Data from the Tritium Systems Test Assembly Program

The Tritium Systems Test Assembly (TSTA) at Los Alamos National Laboratory, operational from 1984 to 2001, was critical in advancing fusion fuel cycle technologies, including tritium storage, gas separation, and pumping. TSTA’s contributions, particularly in safe tritium operations, have influenced subsequent fusion projects. This paper discusses the ongoing effort to digitize and catalog TSTA’s historical data to create a searchable resource for the fusion research community. While the long-term objective is to develop a relational database for structured data management, the project remains in the early phase, with current efforts focused on scanning and indexing physical documents. Initial plans for database implementations are also presented, outlining key considerations for structure, query indexing, and standardization. As digitization progresses, future discussions will refine these implantation details to ensure an efficient and comprehensive system. This initiative aims to preserve critical legacy data, enhance the design of tritium system facilities, and support the next generation of fusion energy research.

42 ENGINEERING

Retrieval Augmented Generation for Robust Cyber Defense

In cybersecurity, the ability to efficiently analyze and respond to vulnerabilities, weaknesses, attack patterns, and threat tactics is critical for effective defense strategies. With the increasing complexity and volume of cybersecurity data, traditional methods of querying and retrieving information are often inadequate. To address this challenge, we implemented Retrieval-Augmented Generation (RAG) systems—CyRAG and GraphCyRAG—that integrate large language models (LLMs) with both structured data from relational databases and knowledge graphs such as Neo4j. CyRAG is designed to handle structured data, focusing on CVE (Common Vulnerabilities and Exposures) and CWE (Common Weakness Enumeration) entities to generate accurate and context-rich responses. In contrast, GraphCyRAG leverages Neo4j knowledge graphs to retrieve interconnected information from CVE, CWE, CAPEC (Common Attack Pattern Enumeration and Classification), and ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) datasets. By utilizing Neo4j’s graph-based framework, GraphCyRAG enables deeper traversal of relationships between vulnerabilities and attack patterns, providing cybersecurity analysts with more comprehensive insights into potential attack vectors and mitigation strategies. Our preliminary results demonstrate that integrating knowledge graphs with RAG significantly enhances both the accuracy and depth of threat analysis, allowing for the retrieval of dynamic, real-time data and the generation of contextually aware responses. This approach helps analysts uncover hidden relationships between cyber entities, predict exploit paths, and prioritize mitigation efforts effectively. The integration of RAG with cybersecurity knowledge graphs represents a significant advancement in cybersecurity threat intelligence, enabling more informed decision-making and stronger defense strategies.

97 MATHEMATICS AND COMPUTING

CLEAP Project: OR-SAGE Analysis for MT, UT, and CO States

The OR-SAGE tool is designed to use industry-accepted practices in screening sites and then employ the proper array of data sources through the considerable computational capabilities of GIS technology available at ORNL. The tool was developed to screen the potential for NPP siting on a national and regional basis. However, because of the tool granularity, it is often focused specifically on the immediate area around user sites of interest. If data center siting parameters can be added to OR-SAGE, the ability to evaluate data center siting on a localized scale will be beneficial.1 More than 60 data sets have been collected and processed by ORNL to develop exclusionary, avoidance, and suitability criteria for screening sites for a variety of power generation types, including nuclear power plants. Available site evaluation parameters include population density, slope, seismic activity, proximity to cooling-water sources, proximity to hazard facilities, avoidance of protected lands and floodplains, susceptibility to landslide hazards, and many others. All siting parameters should be considered as flags to inform siting decisions and should not be used to rule in or rule out any NPP site. Once data center siting parameters are identified, appropriate data sets will be collected and processed. The OR-SAGE process is very versatile. Essentially, OR-SAGE is a visual, relational database. The database partitions the contiguous United States, a total of 720 million hectares (~1.8 billion acres), into 100-m by 100-m (1 hectare or ~2.5 acre) cells. The database is tracking just under 700 million individual land cells. Successive suitability criterion is applied to each cell in the database. User-specified thresholds can be applied to each siting parameter data layer. In this manner, a variety of scenarios can be quickly and thoroughly evaluated. Data can be added and/or revised within OR-SAGE to address user interests. Siting security assessment capability is currently being added to OR-SAGE. Security is expected to be of concern at data centers whether it is collocated with a nuclear power generating technology or not. If data center is collocated with a nuclear power generating source, the security threat attractiveness level of both will likely increase. It will be of additional benefit if a potential data center site is also assessed for security vulnerability.

97 MATHEMATICS AND COMPUTING

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING

NE-COST plug-in: Expanding ACCERT's Capabilities for Life-Cycle Cost Modeling

The Algorithm for the Capital Cost Estimation of Reactor Technologies (ACCERT) is a structured methodology and software tool designed to simplify and standardize cost estimation for nuclear reactor technologies [1]. By utilizing a relational database structure and modular cost estimation algorithms, ACCERT delivers a robust, flexible, and scalable framework for evaluating costs across various reactor types and configurations [2]. The recent integration of the NE-COST plugin further expands ACCERT’s scope by introducing detailed life-cycle cost modeling and probabilistic analysis of uncertainties. This addition enables users to evaluate costs across front-end processes such as uranium enrichment and fabrication, as well as back-end activities including waste disposal and geologic storage. Through Monte Carlo statistical cost simulations, the plugin provides probabilistic insights into cost ranges, offering critical decision-making support for stakeholders including reactor developers, policymakers, and researchers.

Zhou, Jia

Three-dimensional reconstruction of laser-direct-drive inertial confinement fusion hot-spot plasma from x-ray diagnostics on the OMEGA laser facility (invited)

A deep-learning convolutional neural network (CNN) is used to infer, from x-ray images along multiple lines of sight, the low-mode shape of the hot-spot emission of deuterium–tritium (DT) laser-direct-drive cryogenic implosions on OMEGA. The motivation of this approach is to develop a physics-informed 3-D reconstruction technique that can be performed within minutes to facilitate the use of the results to inform changes to the initial target and laser conditions for the subsequent implosion. The CNN is trained on a 3D radiation-hydrodynamic simulation database to relate 2D x-ray images to 3D emissivity at stagnation. The CNN accounts for the lack of an absolute spatial reference and the different bands of photon energies in the x-ray images. While previous works studied the effect of mode-1 asymmetries on implosion performance using nuclear diagnostics, this work focuses on the effect of mode 2 inferred from x-ray diagnostics on implosion performance. A current analysis of 19 DT cryogenic implosions indicates there is an upper limit of ~20% reduction in the neutron yield caused by an ℓ = 2 amplitude for ℓ 2 /ℓ 0 ≤ 0.32. Here, these conclusions are supported by 2D simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Progress Report on SFR Metallic Fuel Data Qualification

This report summarizes the progress of SFR metallic fuel qualification related activities, which are focused on providing quality assurance relevant information applicable to experiments irradiated during the Integral Fast Reactor (IFR) program. An overview of the metallic fuel performance data and the associated databases, including the EBR-II Fuels Irradiation & Physics Database (FIPD), Out-of-Pile Transient Database (OPTD), and TREAT Experimental Relational Database (TREXR) is included. The legacy data in the databases, including as-built, post-irradiation examination (PIE), operating parameters, and out-of-pile experiment post-test data are introduced. The SFR metallic fuel Quality Assurance Program Plan (QAPP) and its implementation to qualify these legacy data is described in detail. Important PIE data QA documents and the specifications of seven types of PIE measurements (contact profilometry, laser profilometry, neutron radiography, gamma scan, fission gas release fission gas chemistry, and metallography) are provided. Examples of the implementation of the QAPP to qualify each of those types of PIE data are provided.

Mo, Kun [Argonne National Laboratory (ANL), Argonn

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab

Generalizability of heat-related health risk associations observed in a large healthcare claims database of patients with commercial health insurance

Extreme ambient heat is unambiguously associated with higher risk of illness and death. The Optum Labs Data Warehouse (OLDW), a database of medical claims from US-based patients with commercial or Medicare Advantage health insurance, has been used to quantify heat-related health impacts. Whether results for the insured sub-population are generalizable to the broader population has to our knowledge not been documented. We sought to address this question, for the US population in California from 2012 to 2019. We examined changes in daily rates of emergency department (ED) encounters and in-patient hospitalization encounters for all-causes, heat-related outcomes, renal disease, mental/behavioral disorders, cardiovascular disease, and respiratory disease. OLDW was the source for health data for insured individuals in California, and health data for the broader population were gathered from the California Department of Health Care Access and Information (HCAI). We defined extreme heat exposure as any day in a group of 2 or more days with maximum temperatures exceeding the county-specific 97.5 th percentile and used a space-time-stratified case–crossover design to assess and compare the impacts of heat on health. Average incidence rates of medical encounters differed by dataset. However, rate ratios for ED encounters were similar across datasets for all causes (ratio of incidence rate ratios (rIRR) = 0.989; 95% confidence interval (CI) = 0.973, 1.011), heat-related causes (rIRR = 1.080; 95% CI = 0.999, 1.168), renal disease (rIRR = 0.963; 95% CI = 0.718, 1.292), and mental health disorders (rIRR = 1.098; 95% CI = 1.004, 1.201). Rate ratios for inpatient encounters were also similar. This work presents evidence that OLDW can continue to be a resource for estimating the health impacts of extreme heat.

54 ENVIRONMENTAL SCIENCES

National Energy Water Treatment and Speciation (NEWTS) Database & Dashboard

The Department of Energy's Office of Fossil Energy & Carbon Management (DOE/FECM) through the National Energy Technology Laboratory (NETL) has launched a free online tool, the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard, which can be utilized by community leaders and water researchers to better understand the composition of energy-related wastewater streams. The NEWTS Database and Dashboard provide public access to difficult-to-access datasets, including the original data sources and the processed data forms for input into aqueous chemistry modeling software. The data provided by the tool will help mitigate environmental risks and identify possible sources of valuable critical minerals (CM). The goal of this ASME Power presentation is to highlight the data and capabilities of this free online-tool for obtaining high quality water datasets in formats that are easy for modeling the treatment and recovery of valuable resources from effluent waste stream associated with energy operations.

Siefert, Nicholas

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES

Open database for GPD analyses

This article summarizes the main ideas behind creating an open database proposed for use in the exploration of generalized parton distributions (GPDs). This lightweight database is well suited for GPD phenomenology and is designed to store both experimental and lattice-QCD data. It can also aid in benchmarking GPD-related developments, such as GPD models. The database utilizes a new data format based on the YAML serialization language, enabling the storage of essential information for modern analyses, such as replica values. It includes interfaces for both Python and C++, allowing straightforward integration with analysis codes.

Burkert, V. D. [Thomas Jefferson National Accelera

Measurement of the 28 Si ⁢(𝑛,𝑛′⁢𝛾) cross section with 𝑛, 𝛾, and correlated 𝑛−𝛾 angular distributions

Silicon has become an unavoidable element in the circuitry central to everyday life. In turn, the interactions of silicon isotopes with neutrons for nuclear physics applications, among other motivations, have become increasingly important to understand. The dominant isotope of silicon, 28 Si, is thus of primary interest for enhanced understanding for neutron transport calculations and related investigations. Unfortunately, the existing measurement database for neutron scattering reactions on 28 Si is minimal, and nuclear data evaluations on this topic have not been updated for decades. This article details new measurements of the 28 Si ⁢(𝑛,𝑛′⁢𝛾) reaction utilizing multiple analysis methods available within the correlated gamma neutron array for scattering (CoGNAC). Specifically, high-precision near-threshold results and high-incident-energy results were obtained using the 𝛾-only and correlated 𝑛−𝛾 techniques. First-ever measurements of the correlated 𝑛−𝛾 angular distribution for particles emitted following population of the first excited state in 28 Si were obtained as well, which provide unique insight into theoretical descriptions of the inelastic neutron scattering reaction mechanism itself and detailed guidance for nuclear reaction models. The results agree well with literature data where they exist, and substantially expand on the current database for neutron reactions on 28 Si .

73 NUCLEAR PHYSICS AND RADIATION PHYSICS