Search NASASearch

SEARCH · Search NASA

Results for “data services”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES

Evolution of the ATLAS TDAQ online software framework towards Phase-II upgrade: Use of Kubernetes as an orchestrator of the ATLAS Event Filter computing farm

The ATLAS experiment at the LHC at CERN continuously evolves its TDAQ system to meet the challenges of new physics goals and technological advancements. As ATLAS prepares for the Phase-II Run 4 of the LHC, significant enhancements in the TDAQ Controls and Configuration (TDAQ-CC) tools have been designed to ensure efficient data collection, processing, and management. This abstract presents the evolution of ATLAS TDAQ-CC system leading up to Phase-II Run 4. As part of the evolution towards Phase-II, Kubernetes has been chosen to orchestrate the Event Filter (EF) farm. By leveraging Kubernetes, ATLAS can dynamically allocate computing resources, scale processing capacity in response to changing data taking conditions and ensure high availability of data processing services. The integration of the Kubernetes with the TDAQ Run Control framework enables perfect synchronisation between the experiment’s data acquisition components and the computing infrastructure. We will discuss the architectural considerations and implementation challenges involved in Kubernetes integration with the ATLAS TDAQ-CC system. We will highlight the benefits of using Kubernetes as an EF farm orchestrator, including improved resource utilization, enhanced fault tolerance, and simplified deployment and management of data processing workflows. In addition, we will report on the extensive testing of Kubernetes that was conducted using a farm of 2500 servers within the experiment data taking environment, demonstrating its scalability and robustness in handling the demands of the ATLAS TDAQ system for Phase-II. The adoption of Kubernetes represents a significant step forward in the evolution of ATLAS TDAQ-CC system, aligning with industry best practices in container orchestration.

Corso Radu, Alina [Univ. of California, Irvine, CA

Reducing Data Center Peak Cooling Demand and Energy Costs with Underground Thermal Energy Storage (UTES)

By recent estimates, data center energy demands are projected to consume between 6.7% and 12% of U.S. annual electricity generation by the year 2028, driven primarily by expanded demands from cloud services, big data analytics, and Artificial Intelligence (AI) (Shehabi et al., 2024). As much as 40% of data center total energy consumption are loads associated with the site infrastructure cooling systems, and these are often highly water consumptive (Aljbour et al., 2024). For energy system planners, this presents significant challenges to meeting and managing the anticipated loads, and especially the peak loads of projected data center deployments. Geothermal technologies offer two unique solutions to these challenges: 1) by serving loads through the deployment of new conventional and/or next-generation geothermal power technologies such as EGS and 2) through an often-overlooked opportunity to reduce data center peak cooling loads. The latter is the focus of this paper which explores Cold Underground Thermal Energy Storage ("Cold UTES") as an emerging industrial-scale geothermal cooling solution. This cooling solution is energy efficient, non-water-consumptive, and utilizes long duration energy storage (LDES) on both diurnal and seasonal time scales. Cold UTES has the potential to also function as a virtual power plant (VPP). The US Department of Energy's Geothermal Technologies Office is supporting R&D to understand the grid and system-wide value, costs, and impacts of deploying this emergent cooling solution at scale.

AI

Introduction: Neuromorphic Materials

The explosive growth in data collection and the need to process it efficiently, as well as the desire to automate increasingly complex tasks in transportation, medical care, manufacturing, security and many other fields have motivated a growing interest in neuromorphic computing. Unlike the binary, transistorbased ON/OFF logic gates and separate logic and memory functionalities employed in digital computing, neuromorphic computing is inspired by animal brains that use interconnected synapses and neurons to perform processing, storage and transmission of information at the same location, while only consuming ~20 W or less of power. Motivated by the brain’s efficiency, adaptability, self-learning and resiliency qualities, neuromorphic computing can be broadly defined as an approach to processing and storing information using hardware and algorithms inspired by models of biological neural systems. Present research in neuromorphic computing encompasses approaches that vary significantly in their degree of neuro-inspiration, from systems that only incorporate features such as asynchronous, event-driven operation or use crossbar arrays of non-volatile memory (NVM) elements to accelerate deep neural networks (DNNs), to designs that embrace the extreme parallelism, sparsity, reconfigurability, adaptability, complexity and stochasticity observed in nervous systems. The term ‘neuromorphic’ computing is often credited to Carver Mead, who in the 1980s investigated Si-based analog electronics to replicate functions of the animal retina. Earlier important advances in this field include the work of Frank Rosenblatt, who proposed the concept of the perceptron, Bernard Widrow, who used this concept to build one of the first analog neural networks, the Adaline and many other researchers (see ref. 6 for an historical perspective on neuromorphic computing). With the recent increase in the use of artificial intelligence and large language models, and rising concerns over the associated energy costs, interest in neuromorphic hardware has expanded rapidly. According to some estimates, driven largely by the drastic growth in the training use of artificial intelligence (AI) models using the current computing architectures, the energy cost of computing is projected to reach the energy supply worldwide by 2045. Furthermore, while this is not a realistic outcome, it means that, if more efficient computing technologies are not developed -- soon -- the world will soon become one where demand for energy and market constraints limit the continued increase of societal access to AI and cloud services from data centers. Data centers used for training and use of these models consume hundreds of terawatt hours of electricity, already past 4% of the US electricity demand.

Circuits

A Framework for Identifying Building Energy Models of Localized Utility Service Areas Using Smart Meter Data

Bottom-up load modeling of buildings offers a versatile approach to simulating baseline demand and scenarios of future technology evolution and adoption at the individual building level. This capability is essential to understanding how future load shapes may change with the adoption of electric equipment and vehicles, particularly as it relates to grid planning and infrastructure investments. Traditionally, grid planning techniques have used historical load data to predict future load and infrastructure needs. However, with the anticipated rise in adoption of electrification technologies such as heat pumps and electric vehicles, historical data become less reliable predictors of the future. By employing ResStock, a high-fidelity building stock modeling tool, we can fine-tune electrification scenarios and aggregate models to represent varying geographic resolutions of the grid system, while considering the underlying features of homes. This may enable a more accurate and responsive approach to anticipate and plan for the evolving landscape of energy demands. We present a new framework that leverages building stock energy modeling to identify building models that align with the load shapes and housing attributes of buildings with AMI data. This approach applies two model layers: (1) a classification step that identifies the presence of air conditioning, electric heating, and electric water heating, and (2) an optimization routine that identifies building energy models aligning with load profile data from advanced metering infrastructure meters. This report demonstrates one approach to deploying this framework, and presents results for three test cases that use both modeled and AMI data to assess performance. For a test case using AMI data in Fort Collins, Colorado, we observed a median monthly electricity load CV-RMSE of 16.6%, and a top ten daily heating and cooling median absolute percent error of 7.7% and 8.3%, respectively. For each AMI meter, we identify a set of potential energy models so that downstream use-cases can account for uncertainty driven by variability of baseline technologies and occupant behavior, which impact the response to electrification and energy efficiency scenarios. Our results indicate that ResStock has potential as a scalable solution for modeling residential energy demand at local grid resolutions. Its performance depends on location-specific factors, underlying building characteristics, and the level of aggregation, offering a path towards more precise and adaptive distribution grid planning for the evolving energy landscape.

24 POWER TRANSMISSION AND DISTRIBUTION

Performance Evaluation of Vertical Federated Machine Learning Against Adversarial Threats on Wide-Area Control System: Preprint

Federated machine learning (FL) is gaining significant popularity to develop cybersecurity solutions in power grids because of its advanced capability to support decentralized data handing at local devices, its privacy preservation, and its low-bandwidth requirement. However, the evolving adversarial machine learning (AML) threats raise significant concerns for the cybersecurity of FL architectures. The FL-based split neural network (SplitNN) achieves high performance through the decentralized training of local neural network models while preserving data privacy across multiple entities. In this paper, we propose a methodology for evaluating the performance of a vertical FLbased anomaly detector against different types of AML attacks, including denial-of-service attacks, adversarial data injection attacks, and replay attacks on the trained local models deployed in the grid network. For a case study, we consider the modified IEEE 13-bus system, and we develop SplitNN-based binary and multiclass classification models to detect, locate, and identify different types of data integrity attacks on the volt-watt control with two pooling layers: maximum pooling and AvgPool. Our experimental results, computed through performance metrics, reveal that the severity of these AML attacks varies with the integrated pooling mechanism, the type of classification model, and the nature of the cyberattack. Further, the AML attacks negatively impacted the prediction time per sample for the pretrained SplitNN during the online testing.

adversarial threats

Navigating Economies of Scale and Multiples for Nuclear-Powered Data Centers and Other Applications with High Service Availability Needs

Nuclear energy is increasingly being considered for such targeted energy applications as data centers in light of their high capacity factors and low carbon emissions. This paper focuses on assessing the tradeoffs between economies of scale versus mass production to identify promising reactor sizes to meet data center demands. A framework is then built using the best cost estimates from the literature to identify ideal reactor power sizes for the needs of the given data center. Results should not be taken to be deterministic but highlight the variability of ideal reactor power output against the required demand. While certain advocates claim that with the gigawatts of clean, firm energy needed, large plants are ideal, others advocate for SMRs that can be deployed in large quantities and reap the benefits from learning effects. The findings of this study showcase that identifying the optimal size for a reactor is likely more nuanced and dependent on the application and its requirements. Overall, the study does show potential economic promise for coupling nuclear reactors to data centers and industrial heat applications under certain key conditions and assumptions.

22 GENERAL STUDIES OF NUCLEAR REACTORS

The ABCs of On-Demand Transit (ODT)

On-demand mobility - also referred to as on-demand transit (ODT) - is a form of public mobility that is flexible with respect to where and when service is provided, and ODT deployments have increased significantly in recent years. Transit agencies are becoming increasingly interested in ODT, and due to differing definitions and various service design and business model options, it can be difficult to learn about the emerging industry. This work provides an overview of the definitions of ODT, recent trends internationally and in the U.S., ODT's benefits and challenges (particularly compared to fixed-route transit), three primary service design options, system costs and funding considerations, and a metrics framework for evaluating ODT systems to ensure continued successful performance. Identified benefits include increased service areas, short ride and wait times, increased user flexibility, potential to reduce energy consumption and emissions through shared trips and smaller, right-sized vehicles, increased safety and comfort through door-to-door service, and rich data streams including granular spatio-temporal data that can be analyzed to continuously improve the service. Challenges include scaling ODT service up as small increases in ridership require additional supply to keep service quality high, serving peak times including keeping low wait times, the lack of fixed schedule being challenging for commuters, integrating ODT services with nearby transit systems, and equity for riders without smartphones who cannot track the vehicle in a mobile app. Finally, an overview of seven ODT case studies (in Texas, Missouri, New York, and Ontario, Canada) performed by NREL and related analysis of travel time, energy and emissions, costs, and equity are presented. Initial key findings include: ODT can be cost- and energy-effective compared to fixed-route transit, ODT serves more people than other transit options, and ODT system deployments can be followed by rapid growth.

24 POWER TRANSMISSION AND DISTRIBUTION

Deployment of inference as a service at the US CMS Tier-2 data centers

Coprocessors, especially GPUs, will be a vital ingredient of data production workflows at the HL-LHC. At CMS, the GPU-as-a-service approach for production workflows is implemented by the SONIC project (Services for Optimized Network Inference on Coprocessors). SONIC provides a mechanism for outsourcing computationally demanding algorithms, such as neural network inference, to remote servers, where requests from multiple clients are intelligently distributed across multiple GPUs by a load-balancing service. This talk highlights the recent progress in deploying SONIC at selected U.S. CMS Tier-2 data centers. Using realistic CMS Run3 data processing workflows, such as those containing transformer-based algorithms, we demonstrate how SONIC is integrated into the production-like environment to enable accelerated inference offloading. We will present developments from both the client and server sides, including production job and data center configurations for NVIDIA and AMD GPUs. We will also present performance scaling benchmarks and discuss the challenges of operating SONIC in CMS production, such as server discovery, GPU saturation, fallback server logic, etc.

Holzman, Burt

Comparisons of the v11.1 Orbiting Carbon Observatory‐2 (OCO‐2) X CO2 Measurements With GGG2020 TCCON

The Orbiting Carbon Observatory 2 (OCO-2) is NASA's first Earth observation satellite mission dedicated to studying the sources and sinks of carbon dioxide (CO 2 ) on a global scale. The observations of reflected sunlight are inverted in a retrieval algorithm to produce estimates of the dry air mole-fractions of CO 2 (X CO2 ). The OCO-2 Level 2 data release, version 11.1 (v11.1) retrievals from the Atmospheric Carbon Observations from Space (ACOS) algorithm, includes significant improvements in the X CO2 data product compared to older OCO-2 data versions. This work compares the v11.1 X CO2 from OCO-2 against X CO2 estimates collected from a global ground-based network known as the Total Carbon Column Observing Network (TCCON), OCO-2's primary validation source. The OCO-2 project provides a version of the Level 2 data product, called “lite” files that include calibrated and bias-corrected XCO2 values, accessible together with all OCO-2 data products through the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). This work shows that OCO-2 X CO2 observations made between September 2014 and December 2023, after quality filtering and the application of an averaging kernel correction, agree well with coincident TCCON data for all OCO-2 observational modes of land (nadir, glint, target) and ocean (glint). The aggregated, bias-corrected, and quality-filtered absolute average bias values are less than or equal to 0.20 parts per million (ppm) globally for all OCO-2 observation modes, where the biases do not indicate a statistically significant time dependence. The land nadir/glint mode has the lowest bias value of −0.03 ± 0.85 ppm.

54 ENVIRONMENTAL SCIENCES

Fusion and Fission Energy and Science Directorate and Information Technology Services Directorate HPC Cluster Reduction, Consolidation, and Savings in Data Center Space, Power, and Cooling

This report evaluates the benefits of decommissioning six legacy FFESD purchased HPC clusters and consolidating services and workloads into a new HPC cluster named HELIOS. The findings demonstrate significant reductions in the data center power and cooling requirements, data center footprint, and operational overhead, while simultaneously increasing computational capacity.

97 MATHEMATICS AND COMPUTING

The 200 Gbps Challenge: Imagining HL-LHC analysis facilities

The IRIS-HEP software institute, as a contributor to the broader HEP Python ecosystem, is developing scalable analysis infrastructure and software tools to address the upcoming HL-LHC computing challenges with new approaches and paradigms, driven by our vision of what HL-LHC analysis will require. The institute uses a "Grand Challenge" format, constructing a series of increasingly large, complex, and realistic exercises to show the vision of HL-LHC analysis. Recently, the focus has been demonstrating the IRIS-HEP analysis infrastructure at scale and evaluating technology readiness for production. As a part of the Analysis Grand Challenge activities, the institute executed a "200 Gbps Challenge", aiming to show sustained data rates into the event processing of multiple analysis pipelines. The challenge integrated teams internal and external to the institute, including operations and facilities, analysis software tools, innovative data delivery and management services, and scalable analysis infrastructure. The challenge showcases the prototypes - including software, services, and facilities - built to process around 200 TB of data in both the CMS NanoAOD and ATLAS PHYSLITE data formats with test pipelines. The teams were able to sustain the 200 Gbps target across multiple pipelines. The pipelines focusing on event rate were able to process at over 30 MHz. These target rates are demanding; the activity revealed considerations for future testing at this scale and changes necessary for physicists to work at this scale in the future. The 200 Gbps Challenge has established a baseline on today's facilities, setting the stage for the next exercise at twice the scale.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Understanding the Costs and Barriers of Residential Electrical Panel and Service Upgrades

Upgrading electrical panels and services in U.S. homes represents a significant barrier to building modernization, imposing considerable costs, delays, and procedural uncertainties on homeowners, contractors, utilities, and building departments. Although cost databases document individual infrastructure activities, they rarely capture the integrated, whole-project perspective required to understand how electrical upgrades are planned, executed, and regulated. Existing research is often geographically constrained, narrowly scoped, or derived from limited samples, leaving significant knowledge gaps unaddressed. This study examines the timelines and costs associated with residential electrical panel and service upgrades using data from a national cross- sectional survey conducted in summer 2025. The survey captured responses from 140 stakeholders across 34 states, including building industry professionals, utility staff, and building department staff. Findings indicate that project duration and cost increase substantially with building size though patterns vary by infrastructure type. Most single-family and small multifamily projects were completed within 30 days, whereas medium and large multifamily buildings show longer and less predictable timelines. Service rating increases modestly extend project duration, but building type and size were the primary determinants of delays. Cost patterns were similar: customer-side equipment and panel replacements scale predictably with building size, while utility-side costs are highly variable, often triggered by threshold-driven infrastructure upgrades such as transformer replacements. Permit fees vary considerably between jurisdictions but generally represent a minor proportion of total project costs. These findings identify electrical upgrades as a significant barrier to residential building costs and highlight the need for improved stakeholder coordination, streamlined permitting, and strategies that reduce upgrade requirements while maintaining safety and code compliance.

Casquero-Modrego, Nuria

WFIP3 - NOAA SHIP site - NREL Ceilometer (Vaisala CL51) / Derived Data

NOAA SHIP ceilometer: netCDF L3 data files have level 3 (L3) data that have gone through the calculation service and contain all the data from the algorithms, including mixing layer height values, and quality index data. L3 default files contain L3 data that use the default preset for a live plot. File naming schema: L3_DEFAULT_ _YYYYMMDDHHMM_ _ .nc Name Description: L3 Identification of the data level DEFAULT Identification of the L3 file type CUSTOM OFFLINE STATION_NUMBER WMO station number, if defined YYYYMMDDHHMM UTC time ParameterKey Identification of the advanced algorithm settings. See the table below for an explanation. FREE_FORMAT File suffix, if defined

17 WIND ENERGY

WFIP3 - CACO site - NREL Ceilometer (Vaisala CL51) / Derived Data

CACO ceilometer: netCDF L3 data files have level 3 (L3) data that have gone through the calculation service and contain all the data from the algorithms, including mixing layer height values, and quality index data. L3 default files contain L3 data that use the default preset for a live plot. File naming schema: L3_DEFAULT_ _YYYYMMDDHHMM_ _ .nc Name Description: L3 Identification of the data level DEFAULT Identification of the L3 file type CUSTOM OFFLINE STATION_NUMBER WMO station number, if defined YYYYMMDDHHMM UTC time ParameterKey Identification of the advanced algorithm settings. See the table below for an explanation. FREE_FORMAT File suffix, if defined

17 WIND ENERGY

Data for "Field-scale Evaluation of Ecosystem Service Benefits of Bioenergy Switchgrass"

Purpose-grown perennial herbaceous species are nonfood crops specifically cultivated for bioenergy production and have the potential to secure bioenergy feedstock resources while enhancing ecosystem services. This study assessed soil greenhouse gas emissions (CO2 and N2O), nitrate (NO3-N) leaching reduction potential, evapotranspiration (ET), and water-use efficiency (WUE) of bioenergy switchgrass ( Panicum virgatum L.) in comparison to corn ( Zea mays L.). The study was conducted on field-scale plots in Urbana, IL, during the 2020–2022 growing seasons. Switchgrass was established in 2020 and urea-fertilized at 56 kg N ha−1 year−1. Corn management followed best management practices for the US Midwest, including no-till and 202 kg N ha−1 year−1 fertilization, applied as urea–ammonium nitrate (32%). Our results showed lower direct N2O emissions in switchgrass compared to corn. Although soil CO2 emissions did not differ significantly during the establishment year, emissions in subsequent years were over 50% higher in switchgrass than in corn, likely due to increased belowground biomass, which was over five times higher in switchgrass. Nitrate-N leaching decreased as the switchgrass stand matured, reaching 80% lower than in corn by the third year. Differences in ET and WUE between corn and switchgrass were not significant; however, results indicate a trend toward reduced WUE in switchgrass under drought, driven by lower aboveground biomass production. Our study demonstrates that switchgrass can be implemented at a commercial scale without negatively impacting the hydrological cycle, while potentially reducing N losses through nitrate-N leaching and soil N2O emissions, and enhancing belowground C storage.

field data