Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Aligning NASA Earth Science Data Stewardship with FAIR Principles: Outcomes, Recommendations, and Future Directions

The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR is being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA's existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data more findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.

Data Product↗

Revolutionizing thermal Management in Next-Generation AI data centers: Challenges and breakthrough innovations

Data centers (DCs) serve as critical infrastructure for powering the growth and evolution of AI. Next-generation AI DCs present unique challenges in thermal management driven by unprecedented computational demands. This paper provides a comprehensive summary of key stakeholder perspectives on technology gaps, infrastructure requirements, test bed needs, emerging opportunities, and preliminary solutions related to thermal management for AI DCs. It establishes six strategic pillars of thermal management for next generation AI DC: reliability, deployability, efficiency, resilience, measurability, and valorization. The discussion spans a range of critical topics, including advanced cooling technologies, thermal strategies for emerging modular and edge DCs, system-level optimization and control frameworks, infrastructure planning and grid integration designs, benchmarking approaches, and pathways for waste heat recovery and reuse. The proposed research, development, and demonstration efforts are aimed at accelerating the deployment of AI DCs while ensuring energy efficiency, reliability, safety, and regulatory compliance.

Wang, Pengtao [ORNL] (ORCID:0000000214713429)↗

Reservoir Storage Capacity Change (ResCap)

Overview Storage capacity is an essential reservoir metric that is directly linked to various water management and energy objectives. Accurate reporting and tracking of change in storage over time is crucial for the safe and reliable operation of the associated dam. While storage information is available for many reservoirs through the National Inventory of Dams, additional details, e.g. water elevation levels as well as changes over time are not included. This dataset contains reservoir storage capacities based on conducted surveys in CONUS. To represent changes in a reservoir’s storage over time, the storage capacity as determined by the first and last conducted survey is listed. The level of detail of surveys can vary greatly and improved with technological advancements. Therefore, the type of survey and year when it was conducted is noted. To ensure a fair comparison of storage capacities, the water elevation level along with the corresponding operation of the dam is reported. Structural changes, e.g. heightening of a dam will have an influence on the storage capacity and are therefore also mentioned. A total of 739 different reservoir storage capacity comparisons are listed, with some reservoirs represented more than once (storage capacity comparison at different water elevation levels). Methodology Data were acquired from USBR reservoir survey reports, TWDB lake survey reports and elevation-area-capacity tables, the RSI Web Portal and the NID (USACE, 2024). Initial storage capacity along with year and type of survey record is compared to the most recent reported storage capacity, survey type and year. Comparison elevation in feet as well as comparison elevation type were either extracted from survey reports (USBR, TWDB) or the Web Portal (RSI) and in some cases cross-referenced with data from other sources (Water Management Data, USACE, Water Data for Texas, TWDB).

Chu, Antonia [ORNL] (ORCID:0009000510540427)↗

Structuring and storing signals with their metadata: practical considerations

This chapter aims to provide a comprehensive overview on structuring signal data and their metadata, highlighting key considerations for optimal management and storage. Specifically, it focuses on: (a) the relevance of data organization; (b) what data to store and what to keep; and (c) data management methods. Thus, this chapter answers where and how to store metadata efficiently. Chapter 3 explains what is considered metadata. Chapters 5 and 6 present how to collect certain metadata through dedicated sensor validation tests (Chapter 5) or algorithmic analysis (Chapter 6).

Nicolaï, Niels↗

Evolution of DUNE’s Production System

The DUNE experiment will start running in 2029 and record 30 PB/year of raw waveforms from Liquid Argon TPCs and photon detectors. The size of individual readouts can range from 100 MB to a typical 8 GB full readout of the detector, and even 100 TB for extended readouts from supernova candidates. These data then need to be cataloged, stored and distributed for processing worldwide. This massive amount of data and a heterogeneous computing environment necessitates a powerful and robust distributed computing infrastructure. In the process of building up that infrastructure, DUNE’s production system has recently undergone an overhaul, in which it has integrated 1) a new workflow management system (justIN) 2) a new data catalog (MetaCat) and 3) a state-of-the-art data management system (Rucio). Simulations of DUNE’s Far Detector and its prototypes ProtoDUNE Horizontal Drift (ProtoDUNE-HD) and ProtoDUNE Vertical Drift (ProtoDUNE-VD), as well as data from ProtoDUNE-HD serve as the first tests of this infrastructure.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Urbanization and malaria have a contextual relationship in endemic areas: A temporal and spatial study in Ghana

In West Africa, malaria is one of the leading causes of disease-induced deaths. Existing studies indicate that as urbanization increases, there is corresponding decrease in malaria prevalence. However, in malaria-endemic areas, the prevalence in some rural areas is sometimes lower than in some peri-urban and urban areas. Therefore, the relationship between the degree of urbanization, the impact of living in urban areas, and the prevalence of malaria remains unclear. This study explores this association in Ghana, using epidemiological data at the district level (2015–2018) and data on health, hygiene, and education. We applied a multilevel model and time series decomposition to understand the epidemiological pattern of malaria in Ghana. Then we classified the districts of Ghana into rural, peri-urban, and urban areas using administratively defined urbanization, total built areas, and built intensity. We converted the prevalence time series into cross-sectional data for each district by extracting features from the data. To predict the determinant most impacting according to the degree of urbanization, we used a cluster-specific random forest. We find that prevalence is impacted by seasonality, but the trend of the seasonal signature is not noticeable in urban and peri-urban areas. While urban districts have a slightly lower prevalence, there are still pockets with higher rates within these regions. These areas of high prevalence are linked to proximity to water bodies and waterways, but the rise in these same variables is not associated with the increase of prevalence in peri-urban areas. The increase in nightlight reflectance in rural areas is associated with an increased prevalence. We conclude that urbanization is not the main factor driving the decline in malaria. However, the data indicate that understanding and managing malaria prevalence in urbanization will necessitate a focus on these contextual factors. Finally, we design an interactive tool, ’malDecision’ that allows data-supported decision-making.

60 APPLIED LIFE SCIENCES↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

United States Nuclear Power Reactor Used Nuclear Fuel Database and Applications

The Unified Database (UDB) within STANDARDS serves as the foundational data infrastructure for managing the United States' spent nuclear fuel inventory of 315,111 discharged assemblies totaling 91,036 metric tons of heavy metal. The database organizes this complex inventory through over 200 interconnected tables structured into eight primary attribute categories, supporting integrated analyses across storage, transportation, and disposal domains. Data enters the UDB through the GC-859 Nuclear Fuel Data Survey, which transitioned to web-based collection in 2023, improving data quality through real-time validation. The UDB enables automated generation of input files for nuclear safety analyses, reducing preparation time from weeks to hours while maintaining traceability. Applications include national inventory reporting, Certificate of Compliance assessments, and facility optimization. The three-tier distribution model balances accessibility with security requirements for federal agencies, national laboratories, and research organizations. The UDB provides essential data infrastructure as spent fuel management transitions from site-specific to integrated national campaigns.

Stefanovic, Peter↗

Summary of Analytical Services for the Hanford Site Radionuclide NESHAP Program

This document is a summary of the point source analytical requirements used to demonstrate compliance for the Department of Energy (DOE) Hanford Site operations with 40 Code of Federal Regulations (CFR) Part 61, “National Emission Standards for Hazardous Air Pollutants,” (NESHAP) Subpart H, “National Emission Standards for Emissions of Radionuclides Other Than Radon From Department of Energy Facilities,” and the Washington Administrative Code (WAC) 246-247, “Radiation Protection – Air Emissions.” This reference collects information from multiple source documents and is not intended to create, supersede, replace or over-ride any existing contractual, DOE, federal or state statutes, regulations, compliance agreements, orders, permits, licenses or other requirements. The requirement source document governs where any difference may exist. The Hanford Mission Integration Solutions (HMIS) Environmental organization has been contracted by DOE to manage and report data collected from the sampling and monitoring of radioactive air emissions point sources, colloquially called stacks. The Environmental organization coordinates the analyses and reporting of samples collected at various facilities across the Hanford Site. These facilities operate approximately 52 stacks that require sampling, monitoring or estimating radioactive air emissions. The stacks are operated by Bechtel National, Inc. (BNI), Central Plateau Cleanup Company (CPCCo), Hanford Tank Waste Operations & Closure (H2C), Hanford Laboratory Management and Integration (HLMI), and Pacific Northwest National Laboratory (PNNL). Stack samples from CPCCo, HLMI and H2C facilities are collected by the operating contractor staff, delivered to HMIS, and then shipped to an offsite contracted laboratory for analyses. The field and laboratory sample data uploaded into the Sample Management and Analytical Results Tracking (SMART) database are used to calculate sample volumes and concentrations. Sample concentrations are evaluated for compliance with federal and state regulations, permits, and license requirements. The SMART database also calculates total curies released for sampled point sources and stacks. Point source effluent concentrations and releases are published annually in publicly available reports. The BNI and PNNL operate several DOE-Hanford Field Office (HFO) stacks subject to the requirements of 40 CFR 61, Subpart H and WAC 246-247. The concentrations, curies released and dose modeling evaluation for these stacks are included in the DOE-HFO annual radionuclide NESHAP report. The sample collection, analyses and emissions estimates for these stacks are outside the scope of HMIS contracted responsibilities and not addressed further in this document.

54 ENVIRONMENTAL SCIENCES↗

Evolution of the ATLAS TDAQ online software framework towards Phase-II upgrade: Use of Kubernetes as an orchestrator of the ATLAS Event Filter computing farm

The ATLAS experiment at the LHC at CERN continuously evolves its TDAQ system to meet the challenges of new physics goals and technological advancements. As ATLAS prepares for the Phase-II Run 4 of the LHC, significant enhancements in the TDAQ Controls and Configuration (TDAQ-CC) tools have been designed to ensure efficient data collection, processing, and management. This abstract presents the evolution of ATLAS TDAQ-CC system leading up to Phase-II Run 4. As part of the evolution towards Phase-II, Kubernetes has been chosen to orchestrate the Event Filter (EF) farm. By leveraging Kubernetes, ATLAS can dynamically allocate computing resources, scale processing capacity in response to changing data taking conditions and ensure high availability of data processing services. The integration of the Kubernetes with the TDAQ Run Control framework enables perfect synchronisation between the experiment’s data acquisition components and the computing infrastructure. We will discuss the architectural considerations and implementation challenges involved in Kubernetes integration with the ATLAS TDAQ-CC system. We will highlight the benefits of using Kubernetes as an EF farm orchestrator, including improved resource utilization, enhanced fault tolerance, and simplified deployment and management of data processing workflows. In addition, we will report on the extensive testing of Kubernetes that was conducted using a farm of 2500 servers within the experiment data taking environment, demonstrating its scalability and robustness in handling the demands of the ATLAS TDAQ system for Phase-II. The adoption of Kubernetes represents a significant step forward in the evolution of ATLAS TDAQ-CC system, aligning with industry best practices in container orchestration.

Corso Radu, Alina [Univ. of California, Irvine, CA↗

Comparison of removal and spatial mark‐resight models for estimating wild pig density

Density estimation is critical to effectively manage invasive species and elucidate areas of highest concern. For wild pigs (Sus scrofa), the ability to estimate density is complicated because of their variable home range sizes and social structure. Common methods for estimating density (e.g., mark-recapture) may be unsuitable in management applications because additional data needs to be collected before and after management. Removal models offer a suitable alternative to estimate density changes following management and can be applied broadly across areas where management of wild pigs is ongoing. We collected wild pig removal and camera trap data from 25 private properties ranging in size from approximately 0.5 km 2 to 95 km 2 across 3 ecoregions in South Carolina, USA, from 2020–2023. We compared factors affecting consistency and precision of property-level density estimates between removal and spatial mark-resight (SMR) models. In general, excluding 1 large outlier, density estimates from removal models were between 0.60 and 15.85 wild pigs/km 2 (median = 5.34) with a median coefficient of variation (CV) of 0.76 and 95% confidence intervals for the CV between 0.70 and 0.94. Similarly, excluding 1 large outlier, density estimates from SMR were between 0.22 and 30.97 wild pigs/km 2 (median = 5.48) with a median CV of 0.39 and 95% confidence intervals for the CV between 0.38 and 1.20. We found the precision of removal models was affected primarily by the number of wild pigs dispatched in the removal period (3 months) and the ecoregion in which they were removed. None of the covariates, including the number of recaptures (a corresponding measure of sample size), influenced precision of the SMR models, although recaptures did influence the density estimates. At the individual property level, density estimates from our 2 estimators were dissimilar from each other in approximately 80% of instances, although none of the covariates we examined influenced dissimilarity. Our results provide unique insight into how sample size affects density estimates using 2 common methods and into novel SMR models that incorporate both marked and unmarked detections. In addition, the density estimates in this study can be used as a reference for wild pig densities in common land cover types throughout the southeastern United States.

60 APPLIED LIFE SCIENCES↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

msdlive-cli-distro

MSD-LIVE, the MultiSector Dynamics – Living, Intuitive, Value-adding, Environment, is a flexible and scalable data and code management system combined with a distributed computational platform that will enable MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and multi-model workflows within a robust Community of Practice. MSD-LIVE will facilitate a new open, collaborative, resource-rich, technology-facilitated, community-driven way of doing MSD research.

Lansing, Carina↗

The 200 Gbps Challenge: Imagining HL-LHC analysis facilities

The IRIS-HEP software institute, as a contributor to the broader HEP Python ecosystem, is developing scalable analysis infrastructure and software tools to address the upcoming HL-LHC computing challenges with new approaches and paradigms, driven by our vision of what HL-LHC analysis will require. The institute uses a "Grand Challenge" format, constructing a series of increasingly large, complex, and realistic exercises to show the vision of HL-LHC analysis. Recently, the focus has been demonstrating the IRIS-HEP analysis infrastructure at scale and evaluating technology readiness for production. As a part of the Analysis Grand Challenge activities, the institute executed a "200 Gbps Challenge", aiming to show sustained data rates into the event processing of multiple analysis pipelines. The challenge integrated teams internal and external to the institute, including operations and facilities, analysis software tools, innovative data delivery and management services, and scalable analysis infrastructure. The challenge showcases the prototypes - including software, services, and facilities - built to process around 200 TB of data in both the CMS NanoAOD and ATLAS PHYSLITE data formats with test pipelines. The teams were able to sustain the 200 Gbps target across multiple pipelines. The pipelines focusing on event rate were able to process at over 30 MHz. These target rates are demanding; the activity revealed considerations for future testing at this scale and changes necessary for physicists to work at this scale in the future. The 200 Gbps Challenge has established a baseline on today's facilities, setting the stage for the next exercise at twice the scale.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Accelerating nuclear-integrated data center pursuits in the USA: SWOT analysis, power-thermal management strategies and demonstration plan

Here, this study explores the increasing interest in leveraging nuclear power to meet the escalating energy demands of data centers in the United States (U.S.) by focusing on key factors that contribute to accelerated deployment. The study highlights the importance of N+1/N+2 power supplies (where N is the required number of units), outlines research and innovations in nuclear-integrated data center thermal management and demonstration plan. It also provides updates about status and costing of various reactor system designs. A summarized strengths, weaknesses, opportunities, and threats (SWOT) analysis shows the potential options for grid connectivity, reactors, and site selection. Suitable site discussions consider land and water availability, grid access, and optical fiber connectivity, and the study presents graded prospects for Department of Energy (DOE) sites with a specific example. Community engagement and partnerships are emphasized, particularly the roles of local government, federal agencies, utilities, and data center industry partners, which are crucial for accelerating deployment, business outreach, and approvals. The study provides actionable insights for stakeholders to accelerate the deployment of nuclear-powered data centers.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗