Search NASASearch

SEARCH · Search NASA

Results for “data requests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La

Inter-Kingdom Viral Interactions

Please cite as : Josué A. Rodríguez-Ramos, Amy E. Zimmerman, Ruonan Wu, Sheryl Bell, Trinidad Alfaro, Kirsten Hofmockel, William C. Nelson. 2025. Inter-Kingdom Viral Interactions. [Data Set] PNNL DataHub. This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the above citations for the data package and associated manuscript. Deciphering viral ecology in soils is challenging due to their high physiochemical and community complexity. To enhance detection of sub-communities of DNA and RNA viruses, we applied fractionation approaches to soils collected across a moisture gradient from a grassland field experiment. Analyses included metagenomics and metatranscriptomics of size-fractionated extracellular viruses (i.e., DNA and RNA viromes), metagenomics of bacteria/archaea- or eukaryote-enriched samples, and whole soil metatranscriptomes with rRNA-depletion or polyadenylation enrichment. While RNA virome and whole soil RNA methods captured similar viral diversity, RNA viromes identified longer, higher-quality genomes. Further, we showed that significantly more DNA viruses were active in higher moisture than lower moisture samples, whereas responses by overall diversity vary by genome type (DNA versus RNA genomes). Finally, we demonstrate the power of fractionation approaches for identifying distinct viral communities that infect unique hosts, which has significant implications for ecological investigations, particularly related to interkingdom interactions.

59 BASIC BIOLOGICAL SCIENCES

Bioenergy Feedstock Library Annual Summary Report 2024

The Bioenergy Feedstock Library (BFL), part of the Biomass Feedstock National User Facility (BFNUF) located at Idaho National Laboratory (INL), is a physical sample repository and a web-accessible electronic database. The BFL stores physical and chemical characteristics of biomass and waste carbon sources for energy use, as well as samples generated from U.S. Department of Energy (DOE) Bioenergy Technologies Office (BETO) and U.S. Department of Agriculture-funded projects. The objective of this Bioenergy Feedstock Library Annual Summary Report for 2024, similar to the 2023 Annual Summary Report , is to focus on the updates to: (1) publicly available analytical data and equipment tracked through the BFNUF, (2) significant increases in the physical samples available for request, (3) sample and data archival progress from recent BETO-funded projects, and (4) publicly available data sets created upon request from BETO, INL projects, or outside entities compared to the previous annual summary reports. This report highlights key statistics and available data and information important for INL, BFL users, academics, and industry.

09 BIOMASS FUELS

CoreMS AutoQC Uploader

The invention is a self contained software utility that is deployed on the computer controlling a mass spectrometer. The purpose of the software is to monitor a given directory for files matching a user specified criteria and automatically upload matching files to a remote server as well as trigger a request that the data be processed by the cloud based CoreMS software

Rabus, Jordan [Pacific Northwest National Laborato

Latency Analysis of the Nexus Digital Twin Framework

Real-time digital catalogs are increasingly relied upon to track metadata and connect disparate data sources for cloud-based data integration efforts. One such tool, Deeplynx Nexus is supporting real-time digital twin efforts through event-driven data integration and time-series queries. Nexus’s usefulness for these applications depends critically on how quickly individual records can be uploaded and downloaded, since delays directly affect the responsiveness of any system built on top of it. However, the actual latency a user should expect from Nexus has not been systematically measured before, particularly for the small, frequent transactions typical of live sensor feeds. Here we show that single-record round-trip latency is 61.1 ms on a local Nexus instance and 391.7 ms on the hosted production infrastructure, a roughly 6.4x difference driven primarily by fixed per-request overhead rather than data volume. This overhead dominates at small scale: comparing single-record and ten-record trials suggests approximately 56 ms of each single-record request is fixed connection and authentication cost rather than data-transfer time, meaning batching even a handful of records is substantially more efficient than transmitting them individually. At large batch sizes, this pattern reverses for uploads, which converge to near parity between local and hosted environments by 25,000-50,000 records, while download latency remains persistently 5.7-6.4x slower on hosted infrastructure even at scale. These results suggest that Nexus deployments intended for real-time digital twin applications should prioritize record batching over single-record transactions, and that download-path optimization on hosted infrastructure offers the largest remaining opportunity to reduce latency at scale. We anticipate these baseline measurements will serve as a reference point for future digital twin projects evaluating whether Nexus’s latency profile meets their real-time requirements, and as a benchmark for tracking the effect of future infrastructure or API changes.

99 - GENERAL AND MISCELLANEOUS

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

TEAMER Technical Support for Optimal Control of an Oscillating Surge Wave Energy Converter (CRADA Final Report)

This project will focus on running experiments that evaluate the benefits of using model predictive control (MPC) to optimize power absorbed by a laboratory-scale oscillating surge wave energy converter (OSWEC). MPC is a promising technique to optimize wave energy converter (WEC) behavior while applying system constraints that can help promote structural integrity and device survivability, but there are few studies that experimentally test this control scheme on WECs. Therefore, the Participant is proposing a series of tests that will assess the benefits of MPC experimentally in response to a variety of sea states. For these tests, the Participant will provide the OSWEC device and Data Acquisition (DAQ) system, and request support from the Contractor to use and operate the wave tank for experiments.

16 TIDAL AND WAVE POWER

FY24 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and data analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and crack formation in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), or,in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or more of: height, color, and 16-bit grayscale values as functions of position in a plane projection) to detect signs of surface corrosion and cracking after being trained on similar data, with the features to be detected. Although the initial scope included screening for broader indicators of corrosion, e.g., pitting, identification of potential cracks was prioritized for the past several years at the request of program leadership. Labeled training data is essential to developing the ML algorithm, and enhancements to data labeling capability have been developed to address this essential precursor to application of ML routines. Efficient labeling is particularly important in view of the large volume of data required to train ML algorithms and the relative rarity of cracks in the ICCWR data set. The updated program will read binary data from either LCM, WAMS or SEM files, interrogate data attributes, facilitate user labeling of data for training ML algorithms, execute ML algorithms, output parameters from trained ML algorithms, report ML model accuracy with respect to labeled data, and generate graphical representations for various analyses. In FY24, hourglass neural networks (HNNs) that were initiated in FY22 were further developed and tested using available LCM data, and their performance was tested against that of the alternative U-Net Neural Network algorithm structure. HNNs along with previously developed Convolutional Neural Networks (CNNs) and Deep Neural Networks (DNNs) comprise a suite of ML tools for identification of cracks in the ICCWR

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING

Driver Identification Dataset

The ORNL Driver Identification Dataset was created to collect and analyze driving behavior data from 50 different drivers. Each driver operated a 2014 Kenworth T270 Class 6 truck around Fort Collins, Colorado while various data sources recorded their driving behavior and vehicle performance. The dataset includes CANbus (Controller Area Network) data, GPS data, inertial measurement data, and biometric data from a heart rate monitor. A cyberattack was executed during each drive, which caused multiple dashboard warning lights to illuminate and set the tachometer and speedometer to zero, regardless of actual speed. The attack was stopped either after one minute or if the driver pulled over. By downloading the dataset, you agree to the following: 1) I will not use or disclose the data for any purpose other than Research as that term is defined in 10 CFR 745.102. 2) I will not, under any circumstances, request or accept private or linking identifiers for the data used. 3) I will not attempt to determine the identity of the individuals associated with the data. 4) I will use appropriate safeguards to prevent the use or disclose of the data for any purpose other than Research.

99 GENERAL AND MISCELLANEOUS

NUTRON NESHAPs Dashboard

My project aims to enhance the Nuclear Material Tracking Application (NUTRON) software by integrating a tool for the National Emission Standards for Hazardous Air Pollutants (NESHAP) emissions calculations and presenting this data on a dashboard. Inefficiencies in the current NESHAP calculation process were addressed, which will result in more timely regulatory compliance efforts. Key improvements include revising the transfer request system, incorporating effective dose calculations, and developing a data visualization dashboard into NUTRON. The project involved creating a wireframe, preparing an Engineering Calculations and Analysis Report (ECAR), stakeholder meetings, and providing supplemental documentation. Key findings indicate that the proposed modifications will streamline the NESHAP calculation process, reduce human error, and provide REC personnel with accurate and timely data for material use determinations and dose estimations. Future work focuses on completing the NUTRON modifications and fully integrating the new features, ensuring a more efficient and reliable system for tracking and reporting nuclear material transfers.

99 - GENERAL AND MISCELLANEOUS

CalCharge CRADA000008852 Master Agreement Amendment 1, Battery Consortium – Proprietary Activities

Lawrence Berkeley National Laboratory (LBNL) is partnering with the California Clean Energy Fund to launch CalCharge, an energy storage innovation accelerator, comprised of emerging and established California companies and related organizations developing battery technologies for the electric/hybrid vehicle transportation, the electric grid and consumer electronics markets. The vision of CalCharge is to accelerate the pace of technology innovation, business growth, and cluster development. Calcharge programs will deliver technology acceleration and technical expertise to the energy storage industry, as well as policy and market development support to strengthen the regional economy. LBNL shall collaborate with CalCharge members on the analysis and testing of battery and energy storage technologies. LBNL’s work will include analysis and testing of external design, examination of materials and components either separately or as a whole, and providing data and observations resulting from each collaboration. LBNL shall maintain and provide access to LBNL specialized facilities for research activities performed by or for CalCharge members. In addition, LBNL will provide expertise for short-term consultation, interpretation of testing data, or to clarify technical obstacles if requested by a Member. Over this time frame, Calcharge partnered with several start-ups to provide analytical resources. Those companies include Halotechnics, ZAF Energy Systems, Volkswagen Group of America, Toyota Motor Corporation, and Ensor Inc.

25 ENERGY STORAGE

EV Charging Infrastructure Energization An Overview of Approaches for Simplifying and Accelerating Timelines to Processing EV Charging Load Service Requests

The United States has seen significant growth in electric vehicle (EV) adoption, leading to increased demand for EV charging infrastructure. Over the past decade, EV charging infrastructure site developers, site hosts, and electric distribution utilities have navigated the process to integrate chargers onto the electric grid. Site developers and site hosts have raised the alarm that the integration process for high-powered EV charging projects does not meet the needs of the EV market for timeliness or cost. High-powered charging stations typically require a load service request or an agreement with the local utility to connect to the grid. The process of energizing a new high-powered charging site can be complex and time-consuming, often taking up to 2 years. This timeline is the result of current utility energization processes having been designed for construction projects that take longer to build (i.e., buildings). The specific challenges stem from various factors, including compartmentalization in application processes, the integration of EV charging process approvals with other distributed energy resources (DERs), and the need to ensure grid reliability. The energization process needs to evolve to meet the growing demand for high-powered EV charging. This white paper compiles information gathered through various conversations with key stakeholders, including utilities, utility regulators, EV charging operators, site developers, and authorities having jurisdiction (AHJ) as well as through an extensive literature review. This document identifies the challenges and provides potential solutions to streamline the process of connecting EV charging infrastructure to the power grid in the United States, serving as a starting point for future conversations around these solutions. The solutions noted in this white paper require collaborative efforts among utilities, regulators, and EV charging infrastructure developers to streamline the grid connection process for EV charging infrastructure. They are broadly organized into four areas: 1. Increase data access and transparency: Develop automated load service request tools, integrate hosting capacity and load service request analyses, incorporate EV adoption forecasts, and provide transparency on the processing queue. 2. Improve energization processes and timing: Create fast-track options based on prescreening criteria, provide flexibility or phased approvals in the load service request/interconnection process, build internal knowledge within utilities about EV charging technologies, and provide standardized workforce training. 3. Promote economic efficiency: Right size distribution components to accurately reflect the load requirements of EV charging infrastructure, make proactive investments in grid infrastructure based on EV adoption forecasts and growth projections, and consider energy equity and environmental justice factors such as equitable access to EV charging when planning infrastructure. 4. Improve grid reliability and resilience: Use load management/power control systems (PCS) at EV charging stations, adopt and implement harmonized standards for communication protocols and information models between the EV charging and grid control infrastructure, and address cybersecurity considerations by implementing robust security measures and standards for EV charging infrastructure—with particular emphasis on clarifying the security requirements for the interface to the grid. The objective of the solutions proposed in this white paper is to accelerate the timeline and decrease costs associated with connecting EV charging infrastructure to the grid. Electric utilities, utility regulators, EV charging infrastructure developers, and site hosts will first need to understand which solutions are available in their service territory, and if warranted, which combination of solutions would support their specific needs. Through the successful implementations of solutions at scale detailed here, industry will demonstrate a new and innovative ecosystem where timely deployment and energization of EV charging infrastructure with greater grid resiliency and reliability is a reality.

24 POWER TRANSMISSION AND DISTRIBUTION

The BTSbot-nearby Discovery of SN 2024jlf: Rapid, Autonomous Follow-up Probes Interaction in an 18.5 Mpc Type IIP Supernova

We present observations of the Type IIP supernova (SN) SN 2024jlf, including spectroscopy beginning just 0.7 days (∼17 hr) after first light. Rapid follow-up was enabled by the new BTSbot-nearby program, which involves autonomously triggering target-of-opportunity requests for new transients in Zwicky Transient Facility data that are coincident with nearby (D < 60 Mpc) galaxies and identified by the BTSbot machine learning model. Early photometry and nondetections shortly prior to first light show that SN 2024jlf initially brightened by >4 mag day −1 , quicker than ∼90% of Type II SNe. Early spectra reveal weak flash ionization features: narrow, short-lived (1.3 < τ[days] < 1.8) emission lines of Hα, He II , and C IV . Assuming a wind velocity of v w = 50 km s −1 , these properties indicate that the red supergiant progenitor exhibited enhanced mass loss in the last year before explosion. We constrain the mass-loss rate to $1{0}^{-4}\lt \dot{M}\,[{M}_{\odot }\,{\mathrm{yr}}^{-1}]\lt 1{0}^{-3}$ by matching observations to model grids from two independent radiative hydrodynamics codes. BTSbot-nearby automation minimizes spectroscopic follow-up latency, enabling the observation of ephemeral early-time phenomena exhibited by transients.

core-collapse supernovae

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)

A change language for ontologies and knowledge graphs

Ontologies and knowledge graphs (KGs) are general-purpose computable representations of some domain, such as human anatomy, and are frequently a crucial part of modern information systems. Most of these structures change over time, incorporating new knowledge or information that was previously missing. Managing these changes is a challenge, both in terms of communicating changes to users and providing mechanisms to make it easier for multiple stakeholders to contribute. To fill that need, we have created KGCL, the Knowledge Graph Change Language (https://github.com/INCATools/kgcl), a standard data model for describing changes to KGs and ontologies at a high level, and an accompanying human-readable Controlled Natural Language (CNL). This language serves two purposes: a curator can use it to request desired changes, and it can also be used to describe changes that have already happened, corresponding to the concepts of “apply patch” and “diff” commonly used for managing changes in text documents and computer programs. Another key feature of KGCL is that descriptions are at a high enough level to be useful and understood by a variety of stakeholders—e.g. ontology edits can be specified by commands like “add synonym ‘arm’ to ‘forelimb’” or “move ‘Parkinson disease’ under ‘neurodegenerative disease’.” We have also built a suite of tools for managing ontology changes. These include an automated agent that integrates with and monitors GitHub ontology repositories and applies any requested changes and a new component in the BioPortal ontology resource that allows users to make change requests directly from within the BioPortal user interface. Overall, the KGCL data model, its CNL, and associated tooling allow for easier management and processing of changes associated with the development of ontologies and KGs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Comparative Analysis of Report-Back of Research Results Strategies for Personal Chemical Exposure Data

Background. Report-back of research results (RBRR) is ethically supported and highly requested by participants yet lacks broadly transferable guidelines for RBRR. Effective RBRR must be responsive to target audience needs and may not be addressed by a ‘one-size-fits-all’ approach. Objective. Within a subset of our 19 studies on RBRR, we had the unique opportunity to carry out a comparative analysis of RBRR strategies across cohorts with similar development and evaluation methods, yet distinct in life stage, geography, number and type of chemicals assessed, and community contexts. Methods. We highlight key outcomes from three environmental health studies: an ongoing New York, NY cohort (Fair Start; n=486) and a Detroit, MI cohort (CLEAR; n=34) assessing exposure to ambient urban pollution during pregnancy, and a longitudinal cohort in Houston, TX (Houston-3H) following Hurricane Harvey (n=312). Focus group and survey data were analyzed to identify lessons learned and explore how RBRR supports understanding of environmental health. Results. Commonalities emerged in RBRR development, design, organization, and data visualization, as well as in how RBRR can contribute to an understanding of health-environment connections. Differences included preferences for individual versus community level findings, as well as distinguishable contextual considerations. For pregnancy cohorts, messaging was framed with cultural sensitivity, and to avoid unintended consequences of parental guilt due to prenatal exposures. In the post-disaster Houston-3H study, participants requested additional transparency regarding sampling design and study rationale. Significance. All RBRR case studies reported chemicals without known regulatory or health guidelines, so results were contextualized within the study population. Participants across cohorts requested multi-study comparisons to better understand their results beyond their communities. While foundational RBRR elements (e.g. plain language, graphic organizers) may supersede cohort-specific differences, RBRR should be personalized to encompass perceptions of health across different life-stage, cultural, and environmental contexts.

Vogel, Taylor J.