Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

United States Nuclear Power Reactor Used Nuclear Fuel Database and Applications

The Unified Database (UDB) within STANDARDS serves as the foundational data infrastructure for managing the United States' spent nuclear fuel inventory of 315,111 discharged assemblies totaling 91,036 metric tons of heavy metal. The database organizes this complex inventory through over 200 interconnected tables structured into eight primary attribute categories, supporting integrated analyses across storage, transportation, and disposal domains. Data enters the UDB through the GC-859 Nuclear Fuel Data Survey, which transitioned to web-based collection in 2023, improving data quality through real-time validation. The UDB enables automated generation of input files for nuclear safety analyses, reducing preparation time from weeks to hours while maintaining traceability. Applications include national inventory reporting, Certificate of Compliance assessments, and facility optimization. The three-tier distribution model balances accessibility with security requirements for federal agencies, national laboratories, and research organizations. The UDB provides essential data infrastructure as spent fuel management transitions from site-specific to integrated national campaigns.

Stefanovic, Peter↗

Summary of Analytical Services for the Hanford Site Radionuclide NESHAP Program

This document is a summary of the point source analytical requirements used to demonstrate compliance for the Department of Energy (DOE) Hanford Site operations with 40 Code of Federal Regulations (CFR) Part 61, “National Emission Standards for Hazardous Air Pollutants,” (NESHAP) Subpart H, “National Emission Standards for Emissions of Radionuclides Other Than Radon From Department of Energy Facilities,” and the Washington Administrative Code (WAC) 246-247, “Radiation Protection – Air Emissions.” This reference collects information from multiple source documents and is not intended to create, supersede, replace or over-ride any existing contractual, DOE, federal or state statutes, regulations, compliance agreements, orders, permits, licenses or other requirements. The requirement source document governs where any difference may exist. The Hanford Mission Integration Solutions (HMIS) Environmental organization has been contracted by DOE to manage and report data collected from the sampling and monitoring of radioactive air emissions point sources, colloquially called stacks. The Environmental organization coordinates the analyses and reporting of samples collected at various facilities across the Hanford Site. These facilities operate approximately 52 stacks that require sampling, monitoring or estimating radioactive air emissions. The stacks are operated by Bechtel National, Inc. (BNI), Central Plateau Cleanup Company (CPCCo), Hanford Tank Waste Operations & Closure (H2C), Hanford Laboratory Management and Integration (HLMI), and Pacific Northwest National Laboratory (PNNL). Stack samples from CPCCo, HLMI and H2C facilities are collected by the operating contractor staff, delivered to HMIS, and then shipped to an offsite contracted laboratory for analyses. The field and laboratory sample data uploaded into the Sample Management and Analytical Results Tracking (SMART) database are used to calculate sample volumes and concentrations. Sample concentrations are evaluated for compliance with federal and state regulations, permits, and license requirements. The SMART database also calculates total curies released for sampled point sources and stacks. Point source effluent concentrations and releases are published annually in publicly available reports. The BNI and PNNL operate several DOE-Hanford Field Office (HFO) stacks subject to the requirements of 40 CFR 61, Subpart H and WAC 246-247. The concentrations, curies released and dose modeling evaluation for these stacks are included in the DOE-HFO annual radionuclide NESHAP report. The sample collection, analyses and emissions estimates for these stacks are outside the scope of HMIS contracted responsibilities and not addressed further in this document.

54 ENVIRONMENTAL SCIENCES↗

Evolution of the ATLAS TDAQ online software framework towards Phase-II upgrade: Use of Kubernetes as an orchestrator of the ATLAS Event Filter computing farm

The ATLAS experiment at the LHC at CERN continuously evolves its TDAQ system to meet the challenges of new physics goals and technological advancements. As ATLAS prepares for the Phase-II Run 4 of the LHC, significant enhancements in the TDAQ Controls and Configuration (TDAQ-CC) tools have been designed to ensure efficient data collection, processing, and management. This abstract presents the evolution of ATLAS TDAQ-CC system leading up to Phase-II Run 4. As part of the evolution towards Phase-II, Kubernetes has been chosen to orchestrate the Event Filter (EF) farm. By leveraging Kubernetes, ATLAS can dynamically allocate computing resources, scale processing capacity in response to changing data taking conditions and ensure high availability of data processing services. The integration of the Kubernetes with the TDAQ Run Control framework enables perfect synchronisation between the experiment’s data acquisition components and the computing infrastructure. We will discuss the architectural considerations and implementation challenges involved in Kubernetes integration with the ATLAS TDAQ-CC system. We will highlight the benefits of using Kubernetes as an EF farm orchestrator, including improved resource utilization, enhanced fault tolerance, and simplified deployment and management of data processing workflows. In addition, we will report on the extensive testing of Kubernetes that was conducted using a farm of 2500 servers within the experiment data taking environment, demonstrating its scalability and robustness in handling the demands of the ATLAS TDAQ system for Phase-II. The adoption of Kubernetes represents a significant step forward in the evolution of ATLAS TDAQ-CC system, aligning with industry best practices in container orchestration.

Corso Radu, Alina [Univ. of California, Irvine, CA↗

Comparison of removal and spatial mark‐resight models for estimating wild pig density

Density estimation is critical to effectively manage invasive species and elucidate areas of highest concern. For wild pigs (Sus scrofa), the ability to estimate density is complicated because of their variable home range sizes and social structure. Common methods for estimating density (e.g., mark-recapture) may be unsuitable in management applications because additional data needs to be collected before and after management. Removal models offer a suitable alternative to estimate density changes following management and can be applied broadly across areas where management of wild pigs is ongoing. We collected wild pig removal and camera trap data from 25 private properties ranging in size from approximately 0.5 km 2 to 95 km 2 across 3 ecoregions in South Carolina, USA, from 2020–2023. We compared factors affecting consistency and precision of property-level density estimates between removal and spatial mark-resight (SMR) models. In general, excluding 1 large outlier, density estimates from removal models were between 0.60 and 15.85 wild pigs/km 2 (median = 5.34) with a median coefficient of variation (CV) of 0.76 and 95% confidence intervals for the CV between 0.70 and 0.94. Similarly, excluding 1 large outlier, density estimates from SMR were between 0.22 and 30.97 wild pigs/km 2 (median = 5.48) with a median CV of 0.39 and 95% confidence intervals for the CV between 0.38 and 1.20. We found the precision of removal models was affected primarily by the number of wild pigs dispatched in the removal period (3 months) and the ecoregion in which they were removed. None of the covariates, including the number of recaptures (a corresponding measure of sample size), influenced precision of the SMR models, although recaptures did influence the density estimates. At the individual property level, density estimates from our 2 estimators were dissimilar from each other in approximately 80% of instances, although none of the covariates we examined influenced dissimilarity. Our results provide unique insight into how sample size affects density estimates using 2 common methods and into novel SMR models that incorporate both marked and unmarked detections. In addition, the density estimates in this study can be used as a reference for wild pig densities in common land cover types throughout the southeastern United States.

60 APPLIED LIFE SCIENCES↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

msdlive-cli-distro

MSD-LIVE, the MultiSector Dynamics – Living, Intuitive, Value-adding, Environment, is a flexible and scalable data and code management system combined with a distributed computational platform that will enable MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and multi-model workflows within a robust Community of Practice. MSD-LIVE will facilitate a new open, collaborative, resource-rich, technology-facilitated, community-driven way of doing MSD research.

Lansing, Carina↗

The 200 Gbps Challenge: Imagining HL-LHC analysis facilities

The IRIS-HEP software institute, as a contributor to the broader HEP Python ecosystem, is developing scalable analysis infrastructure and software tools to address the upcoming HL-LHC computing challenges with new approaches and paradigms, driven by our vision of what HL-LHC analysis will require. The institute uses a "Grand Challenge" format, constructing a series of increasingly large, complex, and realistic exercises to show the vision of HL-LHC analysis. Recently, the focus has been demonstrating the IRIS-HEP analysis infrastructure at scale and evaluating technology readiness for production. As a part of the Analysis Grand Challenge activities, the institute executed a "200 Gbps Challenge", aiming to show sustained data rates into the event processing of multiple analysis pipelines. The challenge integrated teams internal and external to the institute, including operations and facilities, analysis software tools, innovative data delivery and management services, and scalable analysis infrastructure. The challenge showcases the prototypes - including software, services, and facilities - built to process around 200 TB of data in both the CMS NanoAOD and ATLAS PHYSLITE data formats with test pipelines. The teams were able to sustain the 200 Gbps target across multiple pipelines. The pipelines focusing on event rate were able to process at over 30 MHz. These target rates are demanding; the activity revealed considerations for future testing at this scale and changes necessary for physicists to work at this scale in the future. The 200 Gbps Challenge has established a baseline on today's facilities, setting the stage for the next exercise at twice the scale.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Accelerating nuclear-integrated data center pursuits in the USA: SWOT analysis, power-thermal management strategies and demonstration plan

Here, this study explores the increasing interest in leveraging nuclear power to meet the escalating energy demands of data centers in the United States (U.S.) by focusing on key factors that contribute to accelerated deployment. The study highlights the importance of N+1/N+2 power supplies (where N is the required number of units), outlines research and innovations in nuclear-integrated data center thermal management and demonstration plan. It also provides updates about status and costing of various reactor system designs. A summarized strengths, weaknesses, opportunities, and threats (SWOT) analysis shows the potential options for grid connectivity, reactors, and site selection. Suitable site discussions consider land and water availability, grid access, and optical fiber connectivity, and the study presents graded prospects for Department of Energy (DOE) sites with a specific example. Community engagement and partnerships are emphasized, particularly the roles of local government, federal agencies, utilities, and data center industry partners, which are crucial for accelerating deployment, business outreach, and approvals. The study provides actionable insights for stakeholders to accelerate the deployment of nuclear-powered data centers.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

Viskores: Integrating Parallel Scientific Visualization Research into Applications

Viskores is a scientific visualization library that is the primary deployment of such algorithms to the parallel accelerated processors of modern DOE supercomputers. In this paper, we review the capabilities provided by Viskores and how these capabilities are leveraged by other software in the high-performance computing ecosystem. We discuss the Viskores data representation and pay particular attention to array management. Through this array management we describe how data is adapted between Viskores and other software along with strategies for converting dynamic, polymorphic objects to static representations better suited to GPU processing. We conclude with several examples of Viskores integrating with high-performance software that is used in production today.

Moreland, Ken [ORNL] (ORCID:0000000270513288)↗

Adaptable Standards for Discovery, Access, and Usability of Oak Ridge National Laboratory’s Data Portals and Catalogs

Oak Ridge National Laboratory (ORNL) is leveraging its established capabilities and subject matter expertise in data curation, governance, management, national security, and risk assessment and mitigation to support the US Department of Energy (DOE) Grid Modernization Initiative. Using standards modeled by the National Institute of Standards and Technology (NIST), the Data Curation Network (DCN), the Oak Ridge Leadership Computing Facility (OLCF), and other leading organizations in the fields of energy research, high-performance computing, and national and homeland security, ORNL seeks to provide a federated approach to research data discovery, use, and interoperability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Novel and Scalable Method for Microencapsulating Salt Hydrate Phase Change Materials in Core–Shell Fibers

Phase change materials (PCMs) are in high demand for applications such as thermal energy storage in buildings, electronics cooling, and thermal management of electric vehicle batteries and data centers. Among these materials, salt hydrate PCMs are particularly attractive due to their high thermal energy storage capacity and low cost. However, they suffer from two major issues: leakage in the melted phase and phase segregation during phase transitions. Microencapsulation is the primary process capable of addressing both of these challenges. However, there is no reliable or scalable method available for microencapsulating salt hydrate PCMs. As a result, the full potential of salt hydrates for building and data center applications has yet to be realized. In this work, we present an innovative method for the microencapsulation of salt hydrate PCMs using a co‐axial pushing technique. This process creates core–shell fibers, with the salt hydrate as the core and a polymer as the shell. Our approach demonstrates strong potential for scalable microencapsulation of salt hydrate PCMs. In conclusion, achieving scalability could enable their widespread use in applications such as data center cooling, battery thermal management, and building climate control.

Sharma, Jaswinder [Oak Ridge National Laboratory (↗

Data Placement Optimization for ATLAS in a Multi-Tiered Storage System within a Data Center

Scientific experiments and computations, especially in High Energy Physics, are generating and accumulating data at an unprecedented rate. Effectively managing this vast volume of data while ensuring efficient data analysis poses a significant challenge for data centers, which must integrate various storage technologies. This paper proposes addressing this challenge by designing and developing a precise data popularity prediction model utilizing state-of-theart AI/ML techniques. This model is crafted from the analysis of ATLAS data and access patterns. It enables us to migrate infrequently accessed data to more economical storage media, such as tape drives, while storing frequently accessed data on faster yet costlier storage media like HDD or SSD. This strategic approach ensures data is placed optimally into the appropriate storage classes, thereby maximizing storage capacity while minimizing data access latency for end-users. Furthermore, the paper includes a performance evaluation of the prediction model using various key metrics such as F1 score, accuracy, precision and recall. Finally, we present a prototype use case, leveraging real-world file access data to assess the model’s impact on performance.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Data-driven gradient optimization for field emission management in a superconducting radio-frequency linac

Field emission can cause significant problems in superconducting radio-frequency linear accelerators (linacs). When cavity gradients are pushed higher, radiation levels within the linacs may rise exponentially, causing degradation of many nearby systems. This research aims to utilize machine learning with uncertainty quantification to predict radiation levels at multiple locations throughout the linacs and ultimately optimize cavity gradients to reduce field emission-induced radiation while maintaining the total linac energy gain necessary for the experimental physics program. The optimized solutions show over 40% reductions for both neutron and gamma radiation from the standard operational settings. Published by the American Physical Society 2025

43 PARTICLE ACCELERATORS↗

AEPF: Attention-Enabled Point Fusion for 3D Object Detection

Current state-of-the-art (SOTA) LiDAR-only detectors perform well for 3D object detection tasks, but point cloud data are typically sparse and lacks semantic information. Detailed semantic information obtained from camera images can be added with existing LiDAR-based detectors to create a robust 3D detection pipeline. With two different data types, a major challenge in developing multi-modal sensor fusion networks is to achieve effective data fusion while managing computational resources. With separate 2D and 3D feature extraction backbones, feature fusion can become more challenging as these modes generate different gradients, leading to gradient conflicts and suboptimal convergence during network optimization. To this end, we propose a 3D object detection method, Attention-Enabled Point Fusion (AEPF). AEPF uses images and voxelized point cloud data as inputs and estimates the 3D bounding boxes of object locations as outputs. An attention mechanism is introduced to an existing feature fusion strategy to improve 3D detection accuracy and two variants are proposed. These two variants, AEPF-Small and AEPF-Large, address different needs. AEPF-Small, with a lightweight attention module and fewer parameters, offers fast inference. AEPF-Large, with a more complex attention module and increased parameters, provides higher accuracy than baseline models. Experimental results on the KITTI validation set show that AEPF-Small maintains SOTA 3D detection accuracy while inferencing at higher speeds. AEPF-Large achieves mean average precision scores of 91.13, 79.06, and 76.15 for the car class’s easy, medium, and hard targets, respectively, in the KITTI validation set. Results from ablation experiments are also presented to support the choice of model architecture.

Chemistry↗