Search NASA⌕ Search

SEARCH · Search NASA

Results for “HPC training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Hands-On Curriculum for Training in HPC Cluster Deployment and Management

This paper presents the design, methodology, and outcomes of the High-Performance Computing Technologies (HPCT) course, a hands-on training program focused on the system-side of HPC cluster deployment and administration. Delivered as part of the Master in High Performance Computing (MHPC) program, the course introduces students to key concepts in cluster configuration, including networking, software stack provisioning, job scheduling, and monitoring. Initially taught in person, the course was transitioned to an online format during the COVID-19 pandemic. This shift led to the development of openly available instructional material and a flipped-classroom approach that continues to support both in-person and hybrid delivery. All course materials are publicly available at www.hpc.temple.edu/mhpc/hpc-technology/index.html. By documenting the structure, infrastructure, and evolution of HPCT, this paper offers a model for accessible HPC system training that supports workforce development in computational science.

Posada Correa, Fernando [ORNL] (ORCID:000000022565↗

Pretraining Billion-Scale Geospatial Foundational Models on Frontier

As AI workloads increase in scope, generalization capability becomes challenging for small task-specific models and their demand for large amounts of labeled training samples increases. On the contrary, Foundation Models (FMs) are trained with internet-scale unlabeled data via self-supervised learning and have been shown to adapt to various tasks with minimal fine-tuning. Although large FMs have demonstrated significant impact in natural language processing and computer vision, efforts toward FMs for geospatial applications have been restricted to smaller size models, as pretraining larger models requires very large computing resources equipped with state-of-the-art hardware accelerators. Current satellite constellations collect 100+TBs of data a day, resulting in images that are billions of pixels and multimodal in nature. Such geospatial data poses unique challenges opening up new opportunities to develop FMs. We investigate billion scale FMs and HPC training profiles for geospatial applications by pretraining on publicly available data. We studied from end-to-end the performance and impact in the solution by scaling the model size. Our larger 3B parameter size model achieves up to 30% improvement in top1 scene classification accuracy when comparing a 100M parameter model. Moreover, we detail performance experiments on the Frontier supercomputer, America's first exascale system, where we study different model and data parallel approaches using PyTorch's Fully Sharded Data Parallel library. Specifically, we study variants of the Vision Transformer architecture (ViT), conducting performance analysis for ViT models with size up to 15B parameters. By discussing throughput and performance bottlenecks under different parallelism configurations, we offer insights on how to leverage such leadership-class HPC resources when developing large models for geospatial imagery applications.

Tsaris, Aristeidis (aris)↗

Intro to HPC Bootcamp: Engaging New Communities Through Energy Justice Projects

The U.S. Department of Energy (DOE) is a long-standing leader in research and development of high-performance computing (HPC) in the pursuit of science. However, we face daunting challenges in fostering a robust and diverse HPC workforce. Basic HPC is not typically taught at early stages of students' academic careers, and the capacity and knowledge of HPC at many institutions are limited. Even so, such topics are prerequisites for advanced training programs, internships, graduate school, and ultimately for careers in HPC. To help address this challenge, as part of the DOE Exascale Computing Project's Broadening Participation Initiative, we recently launched the Introduction to HPC Training and Workforce Pipeline Program to provide accessible introductory material on HPC, scalable AI, and analytics. We describe the Intro to HPC Bootcamp, an immersive program designed to engage students from underrepresented groups as they learn foundational HPC skills. Here, the program takes a novel approach to HPC training by turning the traditional curriculum upside down. Instead of focusing on technology and its applications, the bootcamp focuses on energy justice to motivate the training of HPC skills through project-based pedagogy and real-life science stories. Additionally, the bootcamp prepares students for internships and future careers at DOE labs. The first bootcamp, hosted by the advanced computing facilities at Argonne, Lawrence Berkeley, and Oak Ridge National Labs and organized by Sustainable Horizons Institute, took place in August 2023.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

PowerGridworld: A Framework for Multi-Agent Reinforcement Learning in Power Systems [SWR-22-07]

NREL's PowerGridworld provides a modular simulation environment for training heterogenous, grid-aware, multi-agent reinforcement learning (RL) policies at scale. The package enables the user to create component gym environments that can be composed into more complex agents. For example, a grid interactive building environment can be created by composing together component environments each encapsulating the building, PV, and battery physics. These multi-component environments can then be combined into multi-agent simulation where each agent's power consumption/injection becomes an input for solving the optimal power flow on a distribution feeder modeled in OpenDSS. Information from OpenDSS, such as bus voltages and line flows, can be included in the agents' observation spaces to enable grid-aware rewards. The default API for the PowerGridworld simulator conforms to RLLib's MultiAgent API and thus enables distributed training using HPC and cloud resources.

Biagioni, David↗

PPO And Friends

PPO and Friends (PPOAF) is a pytorch implementation of proximal policy optimization for single- and multi-agent reinforcement learning (the PPO), along with several optimizations and add-ons (the Friends) to enable efficient MPI-parallelized model training on HPC clusters.

Maguire, AlisterO↗

Scalable training of graph convolutional neural networks for fast and accurate predictions of HOMO-LUMO gap in molecules

Abstract Graph Convolutional Neural Network (GCNN) is a popular class of deep learning (DL) models in material science to predict material properties from the graph representation of molecular structures. Training an accurate and comprehensive GCNN surrogate for molecular design requires large-scale graph datasets and is usually a time-consuming process. Recent advances in GPUs and distributed computing open a path to reduce the computational cost for GCNN training effectively. However, efficient utilization of high performance computing (HPC) resources for training requires simultaneously optimizing large-scale data management and scalable stochastic batched optimization techniques. In this work, we focus on building GCNN models on HPC systems to predict material properties of millions of molecules. We use HydraGNN, our in-house library for large-scale GCNN training, leveraging distributed data parallelism in PyTorch. We use ADIOS, a high-performance data management framework for efficient storage and reading of large molecular graph data. We perform parallel training on two open-source large-scale graph datasets to build a GCNN predictor for an important quantum property known as the HOMO-LUMO gap. We measure the scalability, accuracy, and convergence of our approach on two DOE supercomputers: the Summit supercomputer at the Oak Ridge Leadership Computing Facility (OLCF) and the Perlmutter system at the National Energy Research Scientific Computing Center (NERSC). We present our experimental results with HydraGNN showing (i) reduction of data loading time up to 4.2 times compared with a conventional method and (ii) linear scaling performance for training up to 1024 GPUs on both Summit and Perlmutter.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Shaping the FutureWorkforce: Challenges and Lessons Learned in HPC Education from National Labs and Computing Centers

Workforce training at national laboratories and computing centers is essential and typically falls into two categories: foundational training for newcomers and advanced training for experienced users. Foundational topics—such as version control, build systems, and basic HPC usage—are largely transferable across institutions, while cluster-specific training varies due to differences in hardware, job schedulers, and local workflows. Training on emerging technologies is split between hardware-specific content and broadly applicable programming paradigms. Here, to reduce redundancy and increase impact, national labs, computing centers, and vendors are collaborating through initiatives like the HPC Training Working Group to share best practices, co-develop materials, and broaden outreach. These coordinated efforts aim to make HPC training more accessible, scalable, and consistent across the community.

HPC↗

Effect of Exercise Training and +Gz Acceleration Training on Men

Countermeasures for reduction in work capacity (maximal oxygen uptake and strength) during spaceflight and enhanced orthostatic intolerance during re-entry, landing and egress from the return vehicle are continuing problems. The purpose for this study was to test the hypothesis that passive-acceleration training; supine, interval, exercise plus acceleration training and exercise combined with acceleration training would improve orthostatic tolerance in ambulatory men; and that addition of the aerobic exercise conditioning would not alter this improved tolerance from that of passive-acceleration training. Seven men (24-38 yr) underwent "Passive" training on the Ames human-powered centrifuge (HPC) for 30 min, "Exercise" training on the cycle ergometer with constant +Gz acceleration; and "Combined" exercise training at 40% to 90% of the HPC +Gz(max) exercise level. Maximal supine exercise loads increased significant (P<0.05) by 8.3% (Passive), 12.6% (Exercise), and by 15.4% (Combined) after training, but their post-training maximal oxygen uptakes and maximal heart rates were unchanged. Maximal time to fatigue (endurance) was unchanged with Passive was increased (P<0.05) with Exercise and Combined training. Thus, the exercise in the Exercise and Combined training Phases resulted in greater maximal loads and endurance without effect on maximal oxygen uptake or heart rate. There was a 4% to 6% increase (P<0.05) in all four quadriceps muscle volumes (right and left) after post-Combined training. Resting pre-tilt heart rate was elevated by 12.9% (P<0.05) only after Passive training suggesting that the exercise training attenuated the HR response. Plasma volume (% Delta) was uniformly decreased by 8% to 14% (P<0.05) at tilt-tolerance pre- vs. post-training indicating essentially no effect of training on the level of hypovolemia. Post-training tilt-tolerance time and heart rate were increased (P<0.05) only with Passive training by 37.8% and by 29.1%, respectively. Thus, addition of exercise training appeared to attenuate the increased Passive tilt-tolerance.

Greenleaf, John E.↗

Tandem Predictions for HPC Jobs

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

HPC↗

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗

Oak Ridge Computing Academy: An HPC cluster deployment and management pilot

The High Performance Computing Technologies (HPCT) course is a hands-on High Performance Computing (HPC) cluster deployment and management training program offered as part of the International School for Advanced Studies (SISSA) and the International Center for Theoretical Physics (ICTP) Master in High Performance Computing (MHPC) specialization. Here, this training program introduces students to key concepts in cluster configuration. which include networking, software stack provisioning, job scheduling, and monitoring. The publicly available course materials feature several examples and underlying methods that are broadly applicable to cluster deployment and management. This paper discusses the design of a new workforce development program at the Oak Ridge National Laboratory that is based on HPCT, the Oak Ridge Computing Academy (ORCA). The ORCA pilot program was hosted by the Oak Ridge Leadership Computing Facility (OLCF) in Summer 2025. As a part of this discussion, HPCT and ORCA course contents and infrastructure are outlined, ORCA participant experiences are detailed, and potential opportunities for improvement are discussed.

Education↗

NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

28 NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning: Preprint

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

Best Practices for NERSC Training

The National Energy Research Supercomputing Center (NERSC) at Lawrence Berkeley National Laboratory (LBNL) organizes approximately 20 training events per year for its 8,000 users from 800 projects, who have varying levels of High Performance Computing (HPC) knowledge and familiarity with NERSC's HPC resources. Due to the novel circumstances of the pandemic, NERSC began transforming our traditional smaller-scale, on-site training events to larger-scale, fully virtual sessions in March 2020. We treated this as an opportunity to try new approaches and improve our training best practices. This paper describes the key practices we have developed since the start of this transformation, including considerations for organizing events; collaboration with other HPC centers and the DOE ECP Program to increase reach and impact of events; targeted emails to users to increase attendance; efficient management of user accounts for computational resource access; strategies for preventing Zoombombing; streamlining the publication of professional-quality, closed-captioned videos on the NERSC YouTube channel for accessibility; effective communication channels for Q&A; tailoring training contents to NERSC user needs via close collaboration with vendors and presenters; standardized training procedures and publishing of training materials; and considerations for planning HPC training topics. Additionally, most of these practices will be continued after the pandemic as effective norms for training.

97 MATHEMATICS AND COMPUTING↗

Accelerating Floating-Point Computations with Intel AMX

Intel AMX is a built-in component of recent Intel CPU architectures, first supported by the Intel Sapphire Rapids in 2023, that enables efficient dense matrix multiplications using mixed precision with low-precision data types. The popularity of mixed-precision algorithms has grown recently, primarily due to their use on GPUs to enhance the efficiency of HPC applications, particularly for the training of large language models. The availability of mixed precision on CPUs represents a cost-effective solution for applications where high speed is not critical. This report shows how to use the Intel AMX accelerator through examples in C++ and Python. The examples will focus on mixed-precision floating-point operations obtained by the use of bfloat16 (or BF16) to accelerate code in single precision. We employ a bottom-up methodology, starting from specific register instructions (TMUL operation) to higher-level applications in libraries such as Intel MKL, PyTorch, and TensorFlow, ensuring a comprehensive understanding of the accelerator's potential. Additionally, we provide insights into the expected performance gains when leveraging the accelerator on the Kestrel HPC machine at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Scalable foundation models for numerical simulations on HPC platforms

In recent years, foundation models (FMs) have begun to reshape numerical simulations on high-performance computing (HPC) platforms. These large, pre-trained AI models enable rapid predictions across a broad range of physical domains, including Earth system modeling, fluid dynamics, materials science, as well as complex multi-modal simulations in aerospace engineering and fusion research. By training on diverse datasets, FMs learn intricate relationships and underlying physical behavior while also enabling the quantification of uncertainty in their predictions. This capability allows simulations that once required days of numerical calculation to be completed in minutes (FM inference), supporting real-time design optimization, uncertainty-aware decision making, and more comprehensive exploration of complex scenarios.

AI↗