Search NASA⌕ Search

SEARCH · Search NASA

Results for “energy efficient datacenter”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

LEED: A Lightwave Energy-Efficient Datacenter

The Lightwave Energy-Efficient Datacenter (LEED) program is a disruptive “green-field” approach that provides a quantum leap in the energy efficiency of datacenters. LEED’s fundamental value proposition is that a novel and re-architected optical network—RotorNet— can deliver “more bandwidth per buck” as well as unique system-level attributes that significantly improve overall datacenter energy efficiency and performance. LEED has developed three system-level testbeds. The first testbed uses calibrated hardware and software power measurements to determine server energy efficiency as a function of network bandwidth and workload. These measurements have shown that increasing network communications bandwidth dramatically increases server energy efficiency providing a realistic path to the overall ENLITENED program goal of doubling the number of transactions per joule. The second testbed demonstrates key hardware: a prototype low-loss, high-port count optical “selector switch”. This switch was fabricated, racked, and tested. Measured switch characteristics include loss, bandwidth, crosstalk, switch time, system-level switch time (including the transceivers), and bit error rate. The third testbed demonstrates a fully working and manufactured pinwheel design which dramatically lowers the cost of design, while delivering high switch radix and low reconfiguration times. The LEED project has tied these three novel photonic switch prototypes together with production servers and software through the development of a novel FPGA-based NIC platform called Corundum. Corundum ensures that the packet-switched protocols supported by commodity operating systems and devices can interface with the Rotor switch design. The LEED group has used this combined hardware and software prototype to characterize applications running at a commercially relevant scale. The project has used a combination of enhanced optical modulation amplitude (OMA) modulators, broadband multiplexers and demultiplexers, avalanche photodiodes, and a novel burst-mode receivers to enable the insertion of LEED-developed optical switches without the need for expensive optical amplification. Our modeling has shown that measured LEED-developed device characteristics can achieve link characteristics of 2 pJ/bit including both transceivers and the Rotor switch. In summary, the LEED program has demonstrated a credible and practical path, through novel hardware and software, to realize the program objectives of ENLITENED. The net result will ensure that the United States maintains its strength in the crucial sector of Information Technology, which is vital to both our economic security and our national security.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

There and Back Again: Reimagining Cryogenic Cooling for Scalable Arrays of Dilution Refrigerators for future Quantum Datacenters

While pulse tube cryocoolers enabled the rapid expansion of dilution refrigerator technology over the past two decades, the transition to large-scale quantum systems is now driving a reassessment of the DR’s higher-temperature-stage cooling strategies and how these systems can be effectively scaled in a modular way. Quasi-wet architectures based on centralized cryoplants and forced-flow helium distribution offer compelling advantages in energy efficiency, operational cost, and scalability. With appropriate redundancy, standardized interfaces, and optimized distribution system designs, these architectures will provide a practical and robust path forward for the next generation of quantum computing infrastructure.

Hansen, B. [Fermilab]↗

Congestion Management Solutions for Enhanced Distribution System Operations with Aggregated Distribution Grid Resources Providing Grid Services and Market Participation

Microgrids and other aggregations of distribution grid resources (DGRs) are poised to actively participate in electricity markets and provide essential grid services in the coming years. In fact, DGRs already play such a role through behind-the-meter (BTM) demand response programs and small-scale BTM dispatchable generation initiatives. At the same time, the rapid growth of artificial intelligence (AI) and cryptocurrency datacenters imposes significant, often unpredictable, demands on the power distribution system. Aggregated DGRs can serve as flexible resources that help mitigate these pressures by using available transmission and distribution capacity more efficiently, supporting resource adequacy and other reliability services, and providing bridge strategies while long-term transmission infrastructure is being developed. The impacts this activity will have on distribution networks are not fully understood and could present significant challenges for distribution utilities due to capacity constraints and the need for congestion management. Technical issues include reverse power flow, variability and possible degradation of equipment integrity, voltage violations, and customer power quality concerns. These issues will likely intensify as electricity market operators across the United States implement Federal Energy Regulatory Commission Order 2222 over the next few years.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data Center High-Temperature Liquid Cooling and Heat Reuse Techno-Economic Study: Preprint

Data centers are energy-intensive facilities with growing demands for efficiency and cost-effective operations. Smaller, more distributed edge inference data centers are expected to proliferate as AI applications require low latency closer to the user of AI tools, which presents a growing opportunity to explore the systems implications of liquid cooling on water and energy use. This study analyzes the implementation of high-temperature liquid cooling systems in a prototypical inference 1-MW data center and explores the potential for heat reuse across varying climates with a goal to optimize energy efficiency, reduce capital and operational costs, and identify opportunities for high-performance cooling and water use reduction infrastructure. This analysis evaluated configurations utilizing a peak day hourly sizing and systems performance spreadsheet to evaluate design and operational conditions from which component sizes, installed cost, operational cost, and performance metrics were determined for the Base case and the Elevated case. The techno-economic analysis included heat reuse applications across a range of heat recovery temperatures and heat rejection options. The analysis shows that high-temperature liquid cooling allows for improved energy efficiency, lower water consumption, and lower capital costs compared to traditional cooling approaches. Transitioning to elevated water inlet/outlet temperatures (50 degrees C/60 degrees C) eliminates the need for chillers, cooling towers, and heat recovery equipment in many scenarios across three distinct climate zones. This results in up to 75% capital cost savings for the cooling and heat recovery equipment, and with significantly reduced water consumption, especially in non-heat reuse applications. Heat generated from data centers can also be repurposed for space heating, domestic hot water, and other applications, and is most cost-effective when data center outlet temperatures exceed 55-60 degrees C.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗