Search NASASearch

SEARCH · Search NASA

Results for “datacenter”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis

Datacenter Explorer

An Unreal Engine plugin for visualizing and interacting with hierarchical datacenter infrastructure through JSON configurations.

Greenwood, Scott

There and Back Again: Reimagining Cryogenic Cooling for Scalable Arrays of Dilution Refrigerators for future Quantum Datacenters

While pulse tube cryocoolers enabled the rapid expansion of dilution refrigerator technology over the past two decades, the transition to large-scale quantum systems is now driving a reassessment of the DR’s higher-temperature-stage cooling strategies and how these systems can be effectively scaled in a modular way. Quasi-wet architectures based on centralized cryoplants and forced-flow helium distribution offer compelling advantages in energy efficiency, operational cost, and scalability. With appropriate redundancy, standardized interfaces, and optimized distribution system designs, these architectures will provide a practical and robust path forward for the next generation of quantum computing infrastructure.

Hansen, B. [Fermilab]

Data Center Power Systems: Architectures, Impact on Grid Reliability, Modeling Considerations, and Megawatt-Scale Hardware Testing [Slides]

This slide deck describes typical power systems of large datacenters along with reliability problems to bulk power systems from large-scale integration of datacenters. The slide deck covers the architecture of datacenter power systems, different power electronic converters used inside datacenters, their operation modes, and R&Dopportunities in maintaining grid stability.

24 POWER TRANSMISSION AND DISTRIBUTION

Nodal capacity expansion planning with flexible large-scale load siting

We propose explicitly incorporating large-scale load siting into a stochastic nodal power system capacity expansion planning model that concurrently co-optimizes generation, transmission, and storage expansion. The potential operational flexibility of some of these large loads is also taken into account by considering them as consisting of a set of tranches with different reliability requirements, which are modeled as a constraint on expected served energy across operational scenarios. We implement our model as a two-stage stochastic mixed-integer optimization problem with cross-scenario expectation constraints. To overcome the challenge of scalability, we build upon existing work to implement this model on a high performance computing platform and exploit scenario parallelization using an augmented Progressive Hedging Algorithm. The algorithm is implemented using the bounding features of mpisppy, which have shown to provide satisfactory provable optimality gaps despite the absence of theoretical guarantees of convergence. We test our approach and assess the value of this proactive planning framework on total system cost and reliability metrics using realistic testcases geographically assigned to San Diego and South Carolina, with datacenter and direct air capture facilities as large loads.

24 POWER TRANSMISSION AND DISTRIBUTION

Demonstrating the data center as a flexible grid asset using a C-HIL setup

Increasing data center demand is outpacing grid infrastructure development. Artificial intelligence workloads and hyperscale cloud growth are creating unprecedented demand for power, while traditional grid expansion faces multiyear development timelines. Verrus is developing an innovative datacenter solution for this challenge, data centers that act as active grid-supportive assets rather than passive loads. Our approach integrates a novel grid-aware power flow management system with battery energy storage systems(BESS) into a microgrid-controlled, medium-voltage power distribution architecture that delivers critical capabilities, such as: * Fast response to grid disturbances such over/ under voltage or over/ under frequency * Demand flexibility that can service requests from the utility within 10 s * Uninterrupted transition to islanded operation during grid outages * Continuous uptime assurance for compute loads while maintaining all customer service level agreements. Through Verrus' strategic partnership with the National Renewable Energy Laboratory (NREL), these capabilities were validated using NREL's Advanced Research on Integrated Energy Systems (ARIES) virtual emulation environment to model a 70-MW grid-interactive data center. This paper outlines the design, methodology, and results of this emulated deployment, demonstrating that data centers can provide both critical load resilience and ancillary grid support without compromising uptime requirements. Specifically, we present a digital real time simulation of a 70 MW data center integrated with a physical microgrid controller, and demonstrate the data center response in the event of a grid voltage and frequency event, utility demand response request and utility outage.

24 POWER TRANSMISSION AND DISTRIBUTION

Technoeconomic analysis of hydrogen storage using 1,4-butanediol (BDO)/γ-butyrolactone (GBL) as a stationary backup power system

Liquid organic hydrogen carriers (LOHCs) are compounds that store and release hydrogen in stable forms at high density. While one-way carriers such as methanol and ammonia have gained attention, their economic advantages are often realized from their use as an export product and direct use as a fuel. Liquid carrier materials that can instead be cycled for energy storage have promise for stationary power applications. In particular, the reversible LOHC system 1,4-butanediol (BDO, H 2 -rich) and gamma-butyrolactone (GBL, H 2 -lean) has a lower enthalpy of dehydrogenation compared to conventional cyclic hydrocarbons and can utilize non-precious metal copper-based catalysts. Here, in this study, BDO/GBL system capital and operating expenses are characterized for vapor phase versus liquid phase hydrogenation and dehydrogenation in a 10 MW backup power application corresponding to sizing of Tier 2 datacenters as well as other critical infrastructure such as hospitals. Costs are benchmarked against two incumbent technologies: a well-established methylcyclohexane/toluene carrier system and compressed gas storage. BDO/GBL storage costs are found to differ substantially between operating modes, with liquid phase hydrogenation coupled with liquid phase dehydrogenation leading to the lowest LCOS of $\$$4.58/kg H 2 in the absence of byproduct formation. In this bounding case, LCOS for the BDO/GBL system is lower than for MCH/TOL ($\$$6.97/kg H 2 ) and compressed gas ($\$$8.48/kg H 2 at 170 bara and $\$$12.05/kg H 2 at 350 bara). However, escalating costs of carrier replacement due to byproduct formation (ranging from an added $\$$6–11/kg H 2 ) illustrate the need for highly selective catalysts to ensure BDO/GBL carrier viability.

BDO/GBL

Advancing microelectronics through nanoscale science: A perspective on needs and opportunities from the nanoscale science research centers

Microelectronics are the cornerstone of the modern world, enhancing our daily lives by providing services such as communications and datacenters. These resources are accessible thanks to the continual pursuit of a deeper understanding of the chemical and physical phenomena underlying the materials synthesis approaches and fabrication processes used to create microelectronic components and subsequently the components' responses to electrical, optical, and other stimuli that are utilized within microelectronic systems. Today, further development of microelectronics requires multidisciplinary expertise across scientific disciplines and fields of study—synthesis, materials characterization, nanoscale fabrication, and performance characterization—with focus placed on comprehending the nanoscale forms and features of microelectronic components. The Nanoscale Science Research Centers (NSRCs) are Department of Energy, Office of Science user facilities that support the international scientific community in advancing nanoscale science and technology. As a key component of the U.S. Government's National Nanotechnology Initiative, the NSRCs enable transformative discoveries by providing world-class facilities, expertise, and collaborative opportunities. Here, in this perspective, we showcase a non-exhaustive cross-section of the capabilities housed at and developed by the NSRCs and their user communities to address fundamental synthesis, metrology, fabrication, and performance considerations toward advancing the development of new microelectronics. Finally, we provide a timely outlook on the next major areas of necessary development in nanoscale sciences to continue the innovation of microelectronics into the next generation.

2D materials

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]

A Study of Performance Portability of Low-bit Fused Matrix-Vector Multiplication Kernels in SYCL

Understanding the causes of performance gaps between a portable programming model and a vendor-specific programming model is important for improving performance portability. This paper studies performance portability of low-bit fused general matrix-vector multiplication kernels in SYCL on vendors’ graphics processing units (GPUs). This work introduces the use case, explains the kernel implementations in detail, evaluates the performance of the CUDA, HIP, and SYCL kernels on datacenter, desktop, and laptop GPUs, and investigates the causes of performance gaps. The results show that loop unrolling, kernel dispatch overhead, and sum reduction contribute to the gaps.

Jin, Zheming [ORNL] (ORCID:000000027197780X)

New Energy Infrastructure Outlook [Slides]

This report provides a perspective on energy infrastructure under development in the continental U.S. as of the end of 2025, focusing on those making significant progress toward achieving commercial operation. Infrastructures covered in this report include power plants, electric transmission, natural gas pipelines, liquefied natural gas terminals, and datacenters. Additionally, this report includes a section on stockpiled volumes of coal, natural gas, and petroleum.

24 POWER TRANSMISSION AND DISTRIBUTION

Congestion Management Solutions for Enhanced Distribution System Operations with Aggregated Distribution Grid Resources Providing Grid Services and Market Participation

Microgrids and other aggregations of distribution grid resources (DGRs) are poised to actively participate in electricity markets and provide essential grid services in the coming years. In fact, DGRs already play such a role through behind-the-meter (BTM) demand response programs and small-scale BTM dispatchable generation initiatives. At the same time, the rapid growth of artificial intelligence (AI) and cryptocurrency datacenters imposes significant, often unpredictable, demands on the power distribution system. Aggregated DGRs can serve as flexible resources that help mitigate these pressures by using available transmission and distribution capacity more efficiently, supporting resource adequacy and other reliability services, and providing bridge strategies while long-term transmission infrastructure is being developed. The impacts this activity will have on distribution networks are not fully understood and could present significant challenges for distribution utilities due to capacity constraints and the need for congestion management. Technical issues include reverse power flow, variability and possible degradation of equipment integrity, voltage violations, and customer power quality concerns. These issues will likely intensify as electricity market operators across the United States implement Federal Energy Regulatory Commission Order 2222 over the next few years.

24 POWER TRANSMISSION AND DISTRIBUTION

Job Scheduler-Driven Power Gateway for High Performance Computing

Power gateways in the form of a microgrid can incorporate multiple distributed energy resources (DER) in either grid forming or grid following mode and support high performance computing (HPC) power profiles including the large load-follow requirements observed in multi-user HPC systems. The microgrid’s flexibility to operate in either grid forming or grid following mode and to actively switch between these modes enables baseline power from multiple non-baseline DER while maintaining high power quality metrics for the HPC system. But this enormous flexibility in demand response and time of use shifting is generally programmed independently of any integration with an HPC job scheduler which can better inform the load shaping by the microgrid. While there are many existing approaches where the HPC job scheduler takes in information from the grid to make queue scheduling decisions, this work takes the opposite view and explores a scheduler where the jobs in the queue can directly impact the settings of the grid. Several HPC scheduler strategies are tested where the jobs in the queue directly impact the settings of a microgrid designed for HPC operation which is driving a datacenter with three classes of HPC architectures. The scheduler operation is shown using a microgrid with 64 kW of solar capacity and 320 kWh of battery over a period of 21 days operating with significant low-follow swings, a throttled grid, cloudy conditions, switching between grid following and grid forming modes, and a wide range of battery states-of-charge all while maintaining high quality power metrics. The scheduler provides a mechanism for the job queue to directly impact a power gateway like a microgrid and to improve HPC power outcomes such as maximizing renewable energy usage

microgrid

Microgrid Integration with High Performance Computing Systems for Microreactor Operation

Multiple nuclear microreactor concepts are currently being developed across several sizes and fuel types with high performance computing (HPC) systems anticipated to be end-users of the power. Nuclear microreactors are small in size, portable, produce less than 10 MW electric, operate autonomously, and have a refueling interval of as many as 10 years. However, their load-follow is also generally limited to 10%/minute or worse whereas the power variance in HPC systems easily exceeds this constraint under normal operations. This study explores an approach that requires no load-follow from the microreactor but integrates the HPC system with a microgrid built from commercial-off-the-shelf components. Three typical HPC architectures are explored in the context of microgrid operation in this study. Components of power quality and transient response are empirically measured for five different HPC load-follow response levels using a self-contained mobile datacenter connected to the microgrid capable of integration with a nuclear microreactor.

microgrids

Electromagnetic Transient Modeling of Large Data Centers for Grid-Level Studies

The magnitude and complexity of electricity usage patterns from large data centers are having significant impacts on the operation and dynamics of the power grid; grid operators and planners require a range of specialized data center models to properly evaluate these impacts and specify technical solutions as needed. Towards addressing this need, Pacific Northwest National Laboratory (PNNL) has developed a library of electromagnetic transient (EMT) models for grid-level studies of data centers called the data center model library (DML). This report describes how the DML was created and how it may properly be used. The models present in the DML are generic models; subject matter expertise and additional technical data are needed to modify these models before they can represent any real data center. However, they will significantly reduce the level of effort required to develop site-specific models and can serve as a common starting point to guide industry towards a more refined consensus. Most of the models within DML are dedicated to representing the power electronics interfaces commonly used in modern data centers, such as double-conversion uninterruptible power supplies and single-phase power factor correction converters. These models are intended for use in grid-level studies and are a simplified aggregation of many small components. That said, background material on the physical and electrical design of large data centers is provided as companion material so that users can be aware of many of the details which have been omitted or streamlined as a matter of practical necessity. Additionally, guidance on the application of EMT analysis for data center interconnection studies is provided, which aids users in identifying when the DML is necessary and what sort of additional model development may be necessary for conducting real-world studies.

24 POWER TRANSMISSION AND DISTRIBUTION

Electromagnetic Transient Modeling of Large Data Centers for Grid-Level Studies: Beta Release

The magnitude and complexity of electricity usage patterns from large data centers are having significant impacts on the operation and dynamics of the power grid; grid operators and planners require a range of specialized data center models to properly evaluate these impacts and specify technical solutions as needed. Towards addressing this need, Pacific Northwest National Laboratory (PNNL) has developed a library of electromagnetic transient (EMT) models for grid-level studies of data centers called the data center model library (DML). This report describes how the DML was created and how it may properly be used. This report details the DML’s beta release, completed in July 2026. This is a revision and expansion of the alpha release, which was made available in January 2026 The models present in the DML are generic models; subject matter expertise and additional technical data are needed to modify these models before they can represent any real data center. However, they will significantly reduce the level of effort required to develop site-specific models and can serve as a common starting point to guide industry towards a more refined consensus. Most of the models within DML are dedicated to representing the power electronics interfaces commonly used in modern data centers, such as double-conversion uninterruptible power supplies and single-phase power factor correction converters. These models are intended for use in grid-level studies and are a simplified aggregation of many small components. That said, background material on the physical and electrical design of large data centers is provided as companion material so that users can be aware of many of the details which have been omitted or streamlined as a matter of practical necessity. Additionally, guidance on the application of EMT analysis for data center interconnection studies is provided, which aids users in identifying when the DML is necessary and what sort of additional model development may be necessary for conducting real-world studies.

electromagnetic transients