Search NASA⌕ Search

SEARCH · Search NASA

Results for “liquid cooling system”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Cooling Matters: Benchmarking Large Language Models and Vision-Language Models on Liquid-Cooled Versus Air-Cooled H100 GPU Systems

The unprecedented growth in artificial intelligence (AI) workloads, recently dominated by large language models (LLMs) and vision-language models (VLMs), has intensified power and cooling demands in data centers. This study benchmarks LLMs and VLMs on two HGX nodes, each with 8× NVIDIA H100 graphics processing units (GPUs), using liquid and air cooling. Leveraging GPU Burn, Weights & Biases, and IPMItool, we collect detailed thermal, power, and computation data. Results show that the liquid-cooled systems maintain GPU temperatures between 41-50$^\circ$C, while the air-cooled counterparts fluctuate between 54-72$^\circ$C under load. This thermal stability of liquid-cooled systems yields 17% higher performance (54 TFLOPs/ GPU vs. 46 TFLOPs/GPU), performance-per-watt, reduced energy overhead, and greater system efficiency than the air-cooled counterparts. These findings underscore the energy and sustainability benefits of liquid cooling, offering a compelling path forward for hyperscale data centers seeking to optimize AI infrastructure. https://github.com/iscaas/Cooling-Matters.

Latif, Imran↗

Advancing Sustainability in Data Centers: Evaluation of Hybrid Air/Liquid Cooling Schemes for IT Payload Using Sea Water

Abstract-The growth in cloud computing, Big Data, AI and high-performance computing (HPC) necessitate the deployment of additional data centers (DC's) with high energy demands. The unprecedented increase in the Thermal Design Power (TDP) of the computing chips will require innovative cooling techniques. Furthermore, DC's are increasingly limited in their ability to add powerful GPU servers by power capacity constraints. As cooling energy use accounts for up to 40% of DC energy consumption, creative cooling solutions are urgently needed to allow deployment of additional servers, enhance sustainability and increase energy efficiency of DC's. The information in this study is provided from Start Campus' Sines facility supported by Alfa Laval for the heat exchanger and CO 2 emission calculations. The study evaluates the performance and sustainability impact of various data center cooling strategies including an air-only deployment and a subsequent hybrid air/water cooling solution all utilizing sea water as the cooling source. Here we evaluate scenarios from 3 MW to 15+1 MW of IT load in 3 MW increments which correspond to the size of heat exchangers used in the Start Campus' modular system design. This study also evaluates the CO 2 emissions compared to a conventional chiller system for all the presented scenarios. Results indicate that the effective use of the sea water cooled system combined with liquid cooled systems improve the efficiency of the DC, plays a role in decreasing the CO 2 emissions and supports in achieving sustainability goals.

97 MATHEMATICS AND COMPUTING↗

Providing Thermal Stability for an Exascale Supercomputer: A Case Study of Frontier's Cooling System

High performance computing (HPC) systems frequently produce large dynamic power swings, even under typical operating conditions, that can present a significant challenge for their direct-liquid cooling systems. Further, the primary cooling loops that must remove this waste heat have response times measured in minutes while the underlying HPC component thermal stress is measured in seconds. The per-socket power demand for both compute processing units (CPUs) and graphic processing units ( GPUs) continues to increase with each successive generation while case temperatures are declining. New HPC systems are expected to exacerbate the challenge of these dynamic power swings and the impact on effective and timely cooling systems. This paper describes the cooling and controls system for Oak Ridge National Laboratory’s Frontier Supercomputer, the first sustained exascale system, as a case study for this situation. The cooling and control system for Frontier demonstrates specific success, but with a number of trade-offs and decisions that suggest further design and operating optimizations for the community at large to consider.

42 ENGINEERING↗

Century: Zap Energy’s 100-kW-Scale Repetitive Sheared-Flow-Stabilized Z -Pinch System with Liquid Metal Cooling

Zap Energy is developing the sheared-flow-stabilized (SFS) Z-pinch concept for commercial applications. The SFS Z pinch relies on plasma self-organization, in the sense that plasma dynamics play a critical role in confinement. Using plasma axial current for confinement and compression eliminates the need for external confinement or heating technologies. This compact magnetic confinement technology could, in turn, provide the basis for a cost-effective deuterium-tritium fusion power plant. In addition to a robust experimental program pushing plasma performance towards breakeven conditions, Zap Energy has parallel programs developing power handling systems suitable for future power plants. Technologies under development include high average-power repetitive pulsed power, high duty-cycle cathodes, and liquid metal wall systems. Century is the name of Zap Energy’s first effort to integrate these three components into an operational system capable of firing non-reacting hydrogen SFS Z-pinch plasmas into a liquid-metal-lined container at sustained repetition rates on the order of 0.1 Hz. Here, the pulsed power driver and liquid metal heat exchanger are both designed to sustain input powers of 100 kW. Construction and initial operations with an interim ~10 kW liquid metal heat exchanger are described.

Century↗

Data Center High-Temperature Liquid Cooling and Heat Reuse Techno-Economic Study: Preprint

Data centers are energy-intensive facilities with growing demands for efficiency and cost-effective operations. Smaller, more distributed edge inference data centers are expected to proliferate as AI applications require low latency closer to the user of AI tools, which presents a growing opportunity to explore the systems implications of liquid cooling on water and energy use. This study analyzes the implementation of high-temperature liquid cooling systems in a prototypical inference 1-MW data center and explores the potential for heat reuse across varying climates with a goal to optimize energy efficiency, reduce capital and operational costs, and identify opportunities for high-performance cooling and water use reduction infrastructure. This analysis evaluated configurations utilizing a peak day hourly sizing and systems performance spreadsheet to evaluate design and operational conditions from which component sizes, installed cost, operational cost, and performance metrics were determined for the Base case and the Elevated case. The techno-economic analysis included heat reuse applications across a range of heat recovery temperatures and heat rejection options. The analysis shows that high-temperature liquid cooling allows for improved energy efficiency, lower water consumption, and lower capital costs compared to traditional cooling approaches. Transitioning to elevated water inlet/outlet temperatures (50 degrees C/60 degrees C) eliminates the need for chillers, cooling towers, and heat recovery equipment in many scenarios across three distinct climate zones. This results in up to 75% capital cost savings for the cooling and heat recovery equipment, and with significantly reduced water consumption, especially in non-heat reuse applications. Heat generated from data centers can also be repurposed for space heating, domestic hot water, and other applications, and is most cost-effective when data center outlet temperatures exceed 55-60 degrees C.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Innovative Advanced Hydrogen Mobile Fueler (Final Technical Report)

The US Department of Energy (DOE) funded a project to design, develop, deploy and analyze the economic viability of an innovative Advanced Hydrogen Mobile Fueler (AHMF). As part of the design activity the project team defined specifications based upon vehicle requirements and compliance with specific fueling performance criteria. The AHMF was originally designed to fuel 10-20 fuel cell vehicles (FCV) per day, consistent with the requirements of the H70 fueling category. The AHMF is able to operate without remote power connections, is modular for easy transport and deployment, and can provide expanded daily capacity and multi-day operations using delivered gaseous hydrogen. Toward the end of the project, due to request from the hydrogen industry and some necessary modifications of the AHMF, the system is now modified to fuel heavy duty vehicles. The project was conducted over two primary phases, each including several key tasks, subtasks, and project milestones. Phase 1 involved the design, development, and construction of the AHMF, moving from a conceptual design through to completion of assembly and testing. In addition to reviewing several different approaches, the team chose to take a conventional station design and modify it for mobile fueling. There were several innovative components included in the design. The high pressure storage was the first to achieve a US Department of Energy (DOE) Special Permit (SP 20391) to transport high pressure hydrogen (95 MPa) in a composite cylinder. The second novel system was a liquid nitrogen (LIN) cooling system with a compact heat exchanger. Without this change, the cooling system would not be able to fit into the AHMF. Phase 2 demonstrated the AHMF by fueling fuel cell buses at a fleet in Pomona, CA. The site was chosen because a temporary fueling solution was required while a fixed station was being installed. The existing permits for the fixed station and non-public access also was a determining factor. Over two months, the buses were fueled 320 times with over 5000 kg of hydrogen from tube trailers. Existing shore power was used, and the average electrical efficiency was 0.13 kWh/kg. The average liquid nitrogen (LIN) consumption was 90.68 scf/kg. The fueling and consumption data was provided to the National Renewable Energy Laboratory (NREL) for analysis. An economic analysis was performed and will be provided in a separate report. The project was successful, but not without challenges and lessons learned. The DOT special permit led the way for the use of high pressure, composite cylinders and is being by multiple other systems and applications. The pandemic along with time for the DOT special permit approval delayed the project for years. The team also believes that hydrogen mobile fueling has a use for the industry, especially during this upcoming phase of expansion. However, it does have its limitations due to high cost to build and operate, and the same permitting challenges as a fixed station. A fully capable system at the speeds and pressures of the AHMF may not be necessary for most applications and would help reduce the cost and increase storage capacity. The on-board generator can be easily replaced with shore power or the wide range of power generation solutions in the marketplace. One of the major results of the project was the development of new code language in National Fire Protection Agency (NFPA) 2 and the International Fire Code (IFC) for on-demand mobile fueling. This will provide guidance for Authorities Having Jurisdiction (AHJ) and user on how to permit temporary fueling sites across the nation.

08 HYDROGEN↗

Thermo-Fluid Modeling Framework for Supercomputer Digital Twins: Part 1, Demonstration at Exascale

A thermo-fluid modeling framework is being developed for ExaDigiT---an open-source framework for developing comprehensive digital twins of liquid-cooled supercomputers. The work is being conducted in two parts, and discussion is divided into two companion papers. The work documented in this paper focuses on the development of a cooling system library in Dymola for the Frontier supercomputer at Oak Ridge National Laboratory. The second part, outlined in a companion paper, focuses on a templating structure called Auto-CSM for easily creating model-agnostic, physics-based thermo-fluid cooling system models for liquid-cooled supercomputers using a text-based schema. The cooling model is being developed using primarily the open-source Transient Simulation Framework of Reconfigurable Models (TRANSFORM) library. The library follows the templating architecture developed within the TRANSFORM library for modeling subsystems. A full-system validation was performed to validate a very simple model that is integrated with the system controls, and the results are presented herein.

Kumar, Vineet↗

Machine Learning Analysis of Temperature-Strain Relationships for Structural Health Monitoring of Pipes: Self-powered wireless sensor system for health monitoring of liquid-sodium cooled fast reactors

This report presents machine learning (ML) analysis of temperature-strain relationships for structural health monitoring of nuclear reactor stainless steel (SS) pipes with the strain gauge sensor directly printed on the pipe with a 3D conformal aerosol jet printer. We investigate correlations for two sensor pairs installed on the same SS304 pipe: commercial K-type thermocouple with a printed gold strain gauge (TC3-SG3), and commercial K-type thermocouple with commercial Kyowa strain gauge (TC0-SG0). The temperature ranges for the sensor pairs TC0-SG0 and TC3-SG3 are 20.00°C to 266.37°C and 39.95°C to 219.28°C respectively. ML algorithms in this study include Linear Regression (baseline method), Ridge Regression, Lasso Regression, and Gradient Boosting. Performance evaluation metrics include Root Mean Square Error (RMSE), Mean Square Error (MSE), Mean Absolute Error (MAE), R 2 Score, and Explained Variance. Using advanced feature engineering techniques, we extracted 27 temperature-based features and 30 strategic inclusion features. The best performance was obtained with the Gradient Boosting method, which achieves prediction accuracy of R 2 = 0.9999 and RMSE = 7.69 μStrain for TC0-SG0, and R 2 = 0.9998 and RMSE = 18.03 μStrain for TC3-SG3. While the temperature-strain correlations are weaker for the gauge directly printed on the pipe than for the commercial strain gauge, deployment-ready performance exceeding industry standards is achieved for both sensor pairs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

EVSE Characterization

NextGen Profiles' EVSE characterization efforts explored performance variability in production EVSE through the use of EV emulation equipment and assessed how different operational conditions influence charging behavior. Data were collected at a frequency of 10 Hz from both the EV emulator and EVSE during each charge session and stored in a time-series database for further analysis. As part of the NextGen Profiles project, characterization of high-power EVSE was performed on both conductive and wireless charging infrastructure; however, only conductive charging data are currently included in this repository. This EVSE characterization was performed over a range of DC output currents and voltages, covering both nominal and off-nominal test conditions. This EVSE characterization dataset includes high-power charging data from two types of 350-kW-capable EVSE using liquid-cooled Combined Charging System-1 (CCS1, North American version) cables and connectors. To protect confidentiality, all EVSE metadata are anonymized, and the publicly released datasets are metered at 10-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

Plant Engineers Solar Energy Handbook: Southern California Region

Discussed in order after the introduction are solar components and systems (collectors, storage, service hot water systems, space heating with liquid and air systems, space cooling, heat pumps and controls); computer programs for system optimization; local solar and weather data; a description of buildings and plants in Southern California applying solar technology; current Federal and California solar legislation; standards, codes and performance testing information; a listing of manufacturers, distributors, and professional services available in Southern California region; and information access. Finally, solar design check lists for those engineers who wish to design their own systems. The program for the Solar Workshop for the Plant Engineer, March 30, 1978, Los Angeles, California is included.

14 SOLAR ENERGY↗

DEVAP-EDDR-TES (Simulation framework for a desiccant assisted air conditioning system with heat pump regeneration and energy storage) [SWR-24-66]

This software is a simulation framework that models a load flexible air conditioner system. The system consists of an evaporatively cooled liquid desiccant air conditioner (eLD-AC) subsystem, an electrically driven desiccant regenerator (EDDR) subsystem, and a stratified liquid desiccant storage (SLDS) subsystem. The software can be used to 1) predict the steady-state performance of the system given user-specified convergence criteria; 2) predict the dynamic performance of the entire system over a typical drive cycle operation subjected to user-specified building thermal loads and desired electrical load profile; 3) evaluate the synergy of all three subsystems operating altogether and improve the energy storage control strategy.

Huang, Ransisi↗

Experiments on a vapor compression air conditioner with liquid desiccants for efficient dehumidification

Buildings require air conditioning systems that not only cool and dehumidify supply air but also provide sufficient ventilation to ensure indoor air quality and occupant comfort. However, standard recirculation systems-which introduce about a 10 % to 20 % fraction of outdoor air-often fail to deliver air that is precisely cooled and dry, particularly because 80-90 % of the ventilation cooling load is latent. Mixing humid ventilation air with recirculated indoor air increases the energy and costs required to condition the air to comfortable levels. Dedicated outdoor air systems (DOASs) are designed to handle this latent dominated ventilation load and thus need to have efficient humidity removal. Many cooling cycles can perform this task. Here we describe a liquid desiccant DOAS, which combines a vapor compression cycle and a liquid desiccant absorber and desorber pair. We present its performance at 26 operating conditions and a thermodynamic model which can accurately predict the moisture removal efficiency. The model's performance predictions have a mean percentage error of 2.5 % and a coefficient of variation of the root mean square error of 7.5 %. We also compare the performance of this vapor-compression-coupled liquid desiccant system with a standard vapor compression system with the same components but no liquid desiccant. For the 26 conditions tested in this study, this comparison shows that adding liquid desiccants lowers the required evaporator cooling load by 21 %, allows for 25 % lower compressor volumetric capacity, and 25 % lower electricity use. Future work will leverage this model to quantify the reduction in annual electricity use across different climates, including the need for a standard vapor compression system to reheat the air during some of the year.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Experimental Validation of Thermal Hydraulic Behavior in Sodium Fast Reactors (SFR) with the Thermal Hydraulic Experimental Test Article (THETA)

Thermal stratification and transition to natural circulation pose two of the largest sources of uncertainty in systems-level modeling of liquid metal-cooled fast reactors. As these phenomena typically develop during transient event sequences, licensing-basis events analyzed using systemslevel models may have considerable uncertainties associated with thermal-hydraulic parameters of the system to account for these phenomena. As a result, the validation basis for these phenomena for systems-level codes is insufficient to fully support the wide range of liquid metal fast reactors being developed in the US. Currently, the most viable path for licensing a design is to take significant conservatisms and maintain sufficiently large safety margins to account for this uncertainty.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Thermo-Fluid Modeling Framework for Supercomputing Digital Twins: Part 2, Automated Cooling Models

The development of digital twins for the purpose of improving the energy efficiency of supercomputing facilities is a non-trivial endeavor that is complicated by the difficulty of creating physics-based thermo-fluid cooling system models (CSMs). Within ExaDigit---an open-source framework for liquid-cooled supercomputing digital twins---a thermo-fluid modeling framework is being developed. This effort has been segmented into two with two companion papers describing each portion of the overall effort. Part 1 focuses on the development of a cooling system library in Dymola for the Frontier supercomputer at Oak Ridge National Laboratory {\cite{Kumar2024}. Part 2, this paper, describes an effort to create a template-based auto-generation methodology for CSMs, called \textit{AutoCSM}. In this paper, an overview of the initial AutoCSM architecture and workflow is provided, along with a practical example using the Oak Ridge Leadership Computing Facility's (OLCF) Frontier supercomputer CSM. AutoCSM will (1) improve ExaDigiT's user accessibility by providing a flexible workflow for modularizing the creation of the CSM system and control logic, (2) decrease the development time of CSMs, and (3) standardize the method for incorporating CSMs into the ExaDigiT framework.

Greenwood, Scott↗

FGMS-poster

Idaho National Laboratory (INL) performs post irradiation examination (PIE) of tri-structural isotropic (TRISO)-coated particle fuel to help qualify it for high temperature gas cooled reactors. TRISO fuel compacts are re-irradiated in the Neutron Radiography Reactor (NRAD) to generate the short lived fission products needed for fission product release testing. The Fuel Accident Condition Simulator (FACS) furnace and newly added Screen Neutron Irradiated Fuel for Failure (SNIFF) furnace heat the compacts in helium to temperatures of up to 2,000°C, prompting fission product release—predominantly gaseous xenon and krypton isotopes and condensable products such as cesium—from failed particles. These released isotopes are transported to a fission gas monitoring system (FGMS 1 or FGMS 3), where they accumulate in cryogenic cold traps and are quantified using high-purity germanium (HPGe) detectors. The addition of SNIFF and FGMS 3 increases throughput by enabling simultaneous testing of multiple compacts. Furthermore, automated INL developed software provides continuous, near-real time monitoring of fission product inventories and manages the liquid nitrogen cooling of the traps. These system enhancements improve the efficiency, data quality, and testing capacity of TRISO fuel performance evaluations.

07 - ISOTOPES AND RADIATION SOURCES↗

SAM Finite Volume Method Development Status Update: GCR Application, Restart, and MultiApp

The System Analysis Module (SAM) is being developed as a modern system analysis code for advanced non-light-water-reactor safety analysis under the U.S. DOE NEAMS program. Previous feasibility studies have demonstrated that a staggered-grid finite volume method (SG-FVM), implemented under the MOOSE framework, can deliver more than an order of magnitude speedup over the existing continuous Galerkin finite element method (CG-FEM) solver for liquid-cooled, incompressible but thermally expandable flow systems. This work extends the previous effort to compressible, gas-cooled reactor applications, where pressure couples directly into the mass equation adding additional nonlinearity into the equation system. New code capabilities are implemented for pebble bed high-temperature gas-cooled reactor (PB-HTGR) analysis, including a pebble bed CoreChannel component, built-in pebble bed effective thermal conductivity model and channel-to-channel crossflow model. The capabilities are tested, benchmarked, and demonstrated for problems with increased level of model and physical complexities, including the HTTU effective thermal conductivity test, the SANA passive cooling test, and a demonstration case using the GPBR200 reactor design covering steady-state operation, DLOFC and PLOFC transients. Across all cases, the SG-FVM solver demonstrated strong robustness and efficiency, and the solutions agree well with reference results and data. The finding of this work proves that SG-FVM is a viable and efficient solver pathway for compressible, gas-cooled reactor system analysis in SAM. In addition, work has been done to successfully support SAM-FVM recover/restart code feature that is essential to reactor safety analysis applications, and MultiApp code feature that is essential to multi-scale and multi-physics simulations. In summary, this work continued from previous feasibility studies, and further demonstrated that the SG-FVM will serve as a strong foundation for SAM’s advanced solver algorithm for future deployment.

Zou, Ling↗