Search NASA⌕ Search

SEARCH · Search NASA

Results for “containers, HPC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The HPC Container Experience on the Summit Supercomputer

Containers are seeing widespread use in the world of High Performance Computing, with many HPC Centers either providing their own containerization solution or adopting existing ones like Singularity and Apptainer. The demand for containerization options come from users who want to take advantage of the portability and reproducibility containers can provide, as well as being able to build and use applications that are only distributed in container form or are otherwise unsuited to natively run in an HPC environment. The users served by the Oak Ridge Leadership Computing Facility are no exception. We go over the past and current containerization offerings at the Oak Ridge Leadership Computing Facility, mainly focusing on the Summit supercomputer. We arrive at using a combination of Podman and Singularity to allow users to build and run containers directly on Summit, without requiring external resources or hardware for any step of the process. We look at a couple of projects running on Summit that greatly benefited from being able to use containers on Summit. And we compare benchmarks running natively and in containers on Summit at different scales, observing minimal performance difference and consistent behavior across all tests.

Abraham, Subil↗

High Performance Computing Peak Shaving for Microreactor Operation

There are multiple nuclear microreactors currently under development that are designed to provide autonomous power for as many as ten or more years without refueling and are designed to power high performance computing (HPC) datacenters. But the load-follow speeds for a nuclear microreactor will be much slower than grid power and slower than the power variance typical of a HPC system. HPC datacenters experience peak power load variance driven by several factors ranging from the operation of cooling systems to remove heat from the servers to supporting a wide range of user application workflows and architectures each with different power signatures. One mechanism to support the limited load-follow of a microreactor is peak shaving where an energy storage mechanism is used to shed peak load and reduce significant power variance. This work explores peak electrical load shaving using uninterruptible power supply (UPS) systems designed for HPC support in the context of peak shaving when operating using a nuclear microreactor with a load-follow limited to 10% of load per minute. Using a self contained HPC datacenter complete with stand-alone cooling system and provisioned with an x86 cluster, an ARM cluster, and a graphics processing unit (GPU) cluster, peak shaving for microreactor operation using the UPS battery backup is explored while running two classes of typical HPC user applications. HPC architecture suitability for microreactor operation under this type of peak shaving is examined.

97 MATHEMATICS AND COMPUTING↗

Working with Multiple MPIs: Overcoming ABI Incompatibility

Scientific applications rely on third-party libraries, which may have multiple implementations. While sharing an application programming interface (API), many of these implementations do not have a shared application binary interface (ABI) and require recompiling. Recompiling can be a long and complex process and sometimes not even an option when the application is shipped binary only. ABI incompatibility strikes at the heart of portability, productivity, and performance by (1) impeding application execution across different HPC and Cloud systems; (2) adding developer hours rebuilding an application; and (3) not taking advantage of host-optimized libraries. This tutorial teaches attendees a portable way to address ABI incompatibility in MPI using the Wi4MPI library. Wi4MPI translates the ABI dynamically from the MPI library used to build the application to a different MPI library available at run time. With Wi4MPI, HPC practitioners can break the portability barrier imposed by ABI incompatibility, potentially increase performance, and increase user productivity. The tutorial is broken down in three components: (1) Understanding ABI compatibility in MPI; (2) Translating MPI libraries dynamically; and (3) Applying dynamic translation to key use cases in HPC, including Containers. If you use more than one MPI library or supercomputer, this tutorial is for you.

97 MATHEMATICS AND COMPUTING↗

Productivity frameworks for HPC

Productivity Frameworks for HPC will include container recipes, build recipes, continuous integration scripts, and other software aimed at testing the portability of containerized HPC software across platforms and interconnects. In particular, it tests the utility of bind-mounting at the MPI layer (rather than the underlying fabric layer) to leverage a standardized protocol and avoid various technical debt and vendor lock-in. Since MPIs are often ABI-incompatible, trampolines such as the open-source Wi4MPI will be tested when such cases arise.

Hanford, Nathan↗

LLNL Microbenchmarks

The LLNL Microbenchmarks repository will contain open-source microbenchmarks for HPC system benchmarking activities. The repository will contain the code for various microbenchmarks developed for Request For Proposal (RFP) response evaluation and system acceptance.

Brink, Stephanie [Lawrence Livermore National Labo↗

ellora-spack-gen

The project contains software to analyze the behavior of large language models at code generation tasks. Publicly available information is used to generate Spack package recipes for HPC developers. The software contains orchestration tooling, analysis, and plotting functionality.

Melone, CaetanoN↗

Software Quality Assurance for High Performance Computing Containers

Software containers are a key channel for delivering portable and reproducible scientific software in high performance computing (HPC) environments. HPC environments are different from other types of computing environments primarily due to usage of the message passing interface (MPI) and drivers for specialized hard- ware to enable distributed computing capabilities. This distinction directly impacts how software containers are built for HPC applications and can complicate software quality assurance efforts including portability and performance. This work introduces a strategy for building containers for HPC applications that adopts layering as a mechanism for software quality assurance. The strategy is demonstrated across three different HPC systems, two of them petaflops scale with entirely different interconnect technologies and/or processor chipsets but running the same container. Performance consequences of the containerization strategy are found to be less than 5-14% while still achieving portable and reproducible containers for HPC systems.

97 MATHEMATICS AND COMPUTING↗

ExaWorks software development kit: a robust and scalable collection of interoperable workflows technologies

Scientific discovery increasingly requires executing heterogeneous scientific workflows on high-performance computing (HPC) platforms. Heterogeneous workflows contain different types of tasks (e.g., simulation, analysis, and learning) that need to be mapped, scheduled, and launched on different computing. That requires a software stack that enables users to code their workflows and automate resource management and workflow execution. Currently, there are many workflow technologies with diverse levels of robustness and capabilities, and users face difficult choices of software that can effectively and efficiently support their use cases on HPC machines, especially when considering the latest exascale platforms. We contributed to addressing this issue by developing the ExaWorks Software Development Kit (SDK). The SDK is a curated collection of workflow technologies engineered following current best practices and specifically designed to work on HPC platforms. We present our experience with (1) curating those technologies, (2) integrating them to provide users with new capabilities, (3) developing a continuous integration platform to test the SDK on DOE HPC platforms, (4) designing a dashboard to publish the results of those tests, and (5) devising an innovative documentation platform to help users to use those technologies. Our experience details the requirements and the best practices needed to curate workflow technologies, and it also serves as a blueprint for the capabilities and services that DOE will have to offer to support a variety of scientific heterogeneous workflows on the newly available exascale HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Evaluating integration and performance of containerized climate applications on a Hewlett Packard Enterprise Cray system

Containers have taken over large swaths of cloud computing as the most convenient way of packaging and deploying applications. The features that containers offer for packaging and deploying applications translate to high performance computing (HPC) as well. At The National Oceanic and Atmospheric Administration, containers provide an easy way to build and distribute complex HPC applications, allowing faster collaboration, portability, and experiment computer environment reproducibility amongst the scientific community. The challenge arises when applications rely on message passing interface (MPI). This necessitates investigation into how to properly run these applications with their own unique requirements and produce performance on par with native runs. We investigate the MPI performance for benchmarks and containerized climate models for various containers covering selection of compiler and MPI library combinations from the Cray provided programming environments on the Cray XC supercomputer GAEA. Performance from the benchmarks and the climate models shows that for the most part containerized applications perform on par with the natively built applications when the system optimized Cray MPICH libraries are bound into the container, and the hybrid model containers have poor performance in comparison. We also describe several challenges and our solutions in running these containers, particularly challenges with heterogeneous jobs for the containerized model runs.

Abraham, Subil↗

VTK-m: Visualization for the Exascale Era and Beyond

A recent trend in modern high-performance computing is the increasing use of hybrid architectures, where the vast majority of performance comes from accelerators. Modern accelerators are based on Graphics Processing Units (GPU) that contain many low power cores that in their aggregate provides an extremely high computation rate. Current and future CPU processors are requiring more explicit parallelism as each successive version of the hardware packs in more cores, and technologies like hyperthreading and vector operations require even more parallel processing to leverage each core’s full potential. As an example, the Frontier supercomputer installed at Oak Ridge National Laboratories recently hit a record breaking 1.1 exaflops1 on the LINPACK HPC benchmark [Shoemaker 2022]. The system contains 37632 AMD MI250x GPUs which requires more than half a billion threads to keep the system fully utilized [Khizeran 2022].VTK-m is a toolkit of scientific visualization algorithms for these emerging processor architectures. VTK-m supports the fine-grained concurrency for data analysis and visualization algorithms required to drive extreme scale computing by providing abstract models for data and execution that can be applied to a variety of algorithms across many different processor architectures.

Bolstad, Mark↗

Integrated GW Farm ABM

This Data Repository includes data used for the integrated groundwater- farm ABM model, raw model output from scenario ensemble, and processed outputs that isolate the groundwater storage depletion outcomes for the 35,000 farm cells. Model Inputs: Farm ABM Inputs: This folder contains the input data used by the integrated groundwater - farm ABM modelling script (Python file) used for the high performance computing (HPC) experiments. The sub-folder "data inputs" contains all of the farm attribute data, while the three files in the folder have the hydrogeological data lookup table (NLDAS Cost Curve Attributes.csv), a lookup table (Theis well function table.csv) for the groundwater cost curve function, and the farm indexes and corresponding NLDAS ids for all of the cells run in this experiment (nldas farms subset final.csv). NLDAS Cost curve hydrogeological data: Hydrogeological data aggregated to 1/8 degree resolution and aligned with the NLDAS grid. Parameters include: water depth below ground surface [meters], subsurface porosity [unitless], aquifer depth from ground surface to aquifer bottom [meters], annual average recharge (USGS: mm, Doll: meters), and three different hydraulic conductivity (K) values (meters/day). The three K values represent the mean value from Gleeson et al. (2018), one standard deviation above the mean from Gleeson et al. (2018), and the de Graaf et al. 2020 modifications to certain lithologies. Additional information about these datasets and their processing are documented in the supplement to Yoon et al. 2025 (in review). Output: Raw outputs: This folder contains a .zip file that has model outputs for the entire scenario ensemble. There is one csv for each farm id, using the format "farm farmid cases.csv". The relationship between the farm id and NLDAS id is defined by the "nldas farms subset final.csv" located in the Farm ABM Inputs folder. Each csv has 625 rows, corresponding to 625 combinations of different scenario parameter values. Each row (scenario) represents the outcome of a 100 year simulation. Columns define scenario settings and summary statistics for each scenario. The first four columns define the scenario settings: "hydro ratio," "econ ratio," "K scenario," and "gamma scenario." The hydro and econ ratios are values passed to the modeling script that influence multipliers for other model parameters, as documented in the supplement to Yoon et al. 2025 (in review). The gamma multiplier is a coefficient multiplier applied to the baseline gamma values (values below 1 represent lower unobserved costs compared to baseline, values above 1 represent higher costs). The K scenario names represent K values of: "low": 0.5 m/d, "int 1": 2.5 m/d, "int 2": 10 m/d, "high": 50 m/d, and "gleeson": mean Gleeson K value. "Perc vol depleted" is the fraction of groundwater depleted at the end of the 100 simulation. Processed Output: Derived depletion outcomes from raw outputs: All of the individual csv files from the Raw outputs were aggregated into a single file that has the scenario settings and fraction depletion "Perc vol depleted" for every farm cell, for every scenario. The other two files define relationships between the farm id, NLDAS id, and local and major aquifer units, used for aquifer-level depletion analysis.

Agent based modeling↗

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗

HPC Center of the Future: R&D Acquisition Intent

This document contains intended technical requirements for an anticipated future procurement for Lawrence Livermore National Laboratory (LLNL), hereafter referred to as “LLNL” or “the Laboratory”, which seeks to fund a few new and innovative proposals for developing concepts and performing studies that would lead to products available in the 2029-2030 timeframe. Our interest is to incentivize explorations and evaluations of technical capabilities in support of our vision for a High Performance Computing (HPC) center of the future at Livermore Computing (LC), the HPC center at LLNL. To realize this vision a new approach to our approach to acquisition is being considered by LC. The approach is designed to decrease risks for the vendors, while encouraging a broader set of vendors able and we hope willing to partner with LLNL and LC. The future HPC Center vision and acquisition approach was conceived to meet the future mission needs of the Advanced Simulation and Computing (ASC) Program within the National Nuclear Security Administration (NNSA).

97 MATHEMATICS AND COMPUTING↗

Electra: A Modular-Based Expansion of NASA's Supercomputing Capability

NASA has increasingly relied on high-performance computing (HPC) re- sources for computational modeling, simulation, and data analysis to meet the science and engineering goals of its missions in space exploration, aeronautics, and Earth and space science. The NASA Advanced Supercomputing (NAS) Division at Ames Research Center in Silicon Valley, Calif., hosts NASA’s premier supercomputing resources, integral to achieving and enhancing the success of the agency’s missions. NAS provides a balanced environment, funded under the High-End Computing Capability (HECC) project, comprised of world-class supercomputers, including its flagship distributed-memory cluster, Pleiades; high-speed networking; and massive data storage facilities, along with multi-disciplinary support teams for user support, code porting and optimization, and large-scale data analysis and scientific visualization. However, as scientists have increased the fidelity of their simulations and engineers are conducting larger parameter-space studies, the requirements for supercomputing resources have been growing by leaps and bounds. With the facility housing the HECC systems reaching its power and cooling capacity, NAS undertook a prototype project to investigate an alternative approach for housing supercomputers. Modular supercomputing, or container-based computing, is an innovative concept for expanding NASA’s HPC capabilities. With modular supercomputing, additional containers—similar to portable storage pods—can be connected together as needed to accommodate the agency’s ever-increasing demand for computing resources. In addition, taking advantage of the local weather permits the use of cooling technologies that would additionally save energy and reduce annual water usage. The first stage of NASA’s Modular Supercomputing Facility (MSF) prototype, which resulted in a 1,000 square-foot module on a concrete pad with room for 16 compute racks, was completed in Fall 2016 and an SGI (now HPE) computer system, named Electra, was deployed there in early 2017. Cooling is performed via an evaporative system built into the module, and preliminary experience shows a Power Usage Effectiveness (PUE) measurement of 1.03. Electra achieved over a petaflop on the LINPACK benchmark, sufficient to rank number 96 on the November 2016 TOP500 list [14]. The system consists of 1,152 InfiniBand-connected Intel Xeon Broadwell-based nodes. Its users access their files on a facility-wide file system shared by all HECC compute assets via Mellanox MetroX InfiniBand extenders, which connect the Electra fabric to Lustre routers in the primary facility over fiber-optic links about 900 feet long. The MSF prototype has exceeded expectations and is serving as a blueprint for future expansions. In the remainder of this chapter, we detail how modular data center technology can be used to expand an existing compute resource. We begin by describing NASA’s requirements for supercomputing and how resources were provided prior to the integration of the Electra module-based system.

Biswas, Rupak↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

Microgrid Integration with High Performance Computing Systems for Microreactor Operation

Multiple nuclear microreactor concepts are currently being developed across several sizes and fuel types with high performance computing (HPC) systems anticipated to be end-users of the power. Nuclear microreactors are small in size, portable, produce less than 10 MW electric, operate autonomously, and have a refueling interval of as many as 10 years. However, their load-follow is also generally limited to 10%/minute or worse whereas the power variance in HPC systems easily exceeds this constraint under normal operations. This study explores an approach that requires no load-follow from the microreactor but integrates the HPC system with a microgrid built from commercial-off-the-shelf components. Three typical HPC architectures are explored in the context of microgrid operation in this study. Components of power quality and transient response are empirically measured for five different HPC load-follow response levels using a self-contained mobile datacenter connected to the microgrid capable of integration with a nuclear microreactor.

microgrids↗

Regen: An object layout regenerator on large-scale production HPC systems

This article proposes an object layout regenerator called Regen which regenerates and removes the object layout dynamically to improve the read performance of applications. Regen first detects frequent access patterns from the I/O requests of the applications. Second, Regen reorganizes the objects and regenerates or preallocates new object layouts according to the identified access patterns. Finally, Regen removes or reuses the obsolete or regenerated object layouts as necessary. As a result, Regen accelerates access to objects by providing a flexible object layout. We implement Regen as a framework on top of Proactive Data Container (PDC) and evaluate it on Cori supercomputer, a production-scale HPC system, by using realistic HPC I/O benchmarks. The experimental results show that Regen improves the I/O performance by up to 16.92 × compared with an existing system.

Distributed file system↗

Virtual Engineering Software Framework for Integrated Biomass Conversion Modeling

This presentation covers the design and implementation of a software tool to systematically connect computational models of unit operations to simulate an integrated process of low-temperature conversion of biomass to fuel. This virtual engineering (VE) software was designed with the overarching goal of connecting unit models written in various programming languages and requiring different computational resources within a single, flexible framework. The models and features currently considered for the VE library include mechanistic models for pretreatment, enzymatic hydrolysis, and aerobic bioreaction; high-fidelity computational fluid dynamics (CFD) simulations for enzymatic hydrolysis and aerobic bioreaction; and the capability to perform techno-economic analyses (TEA) using Aspen Plus, a commercial software package. The CFD models require access to high-performance computing (HPC) resources, so in addition to handling multiple programming languages and interfaces, the VE software must also be capable of interacting with an HPC scheduler to submit, run, and post-process jobs. Using the Python programming language, a new VE software package has been developed that contains functionality to manage the input-output communication between various unit models, schedule simulations to run on NREL's HPC and analyze those results, and interface with existing TEA software workflows. A Jupyter-notebook GUI was also created to solicit user input and provide documentation. In cases where multiple models for a particular unit-operation exist, selection between models is accomplished through a simple checkbox, with the appropriate inputs and outputs being parsed and converted seamlessly in the background. Each operation makes use of a different programming language, but the flow of information from pretreatment to enzymatic hydrolysis to bioreaction is managed with an intuitive, centralized file-communication strategy. In this talk, the programming approach and implementation details of the notebook are presented for multiple possibilities of the conversion process, including a demonstration of the ability to manage HPC resources. Additionally, an example of a sensitivity study of treatment parameters governing the overall conversion outcome is shown which highlights the ease of defining new problems using the VE Notebook workflow and leads into a discussion of ongoing work to enable outer-loop optimization studies.

biofuel↗