Search NASA⌕ Search

SEARCH · Search NASA

Results for “cluster scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing

Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. Furthermore, the experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.

97 MATHEMATICS AND COMPUTING↗

Correlating and Simulating Socio-Demographically Driven Residential End-Use Activity Schedules

Incorporating socio-demographic and behavioral considerations into decision-support tools is crucial for identifying gaps and addressing consumer needs to ensure reliable and affordable energy solutions. In energy simulation models, the correlation between socio-demographics and time-use behavior is not well-captured. Thus, we developed a large-scale simulation workflow to generate schedules for 10 residential activities across 24 population segments defined by age, income, and employment status. Using pre-pandemic 2015-2019 American Time Use Survey (ATUS) data, we used ANOVA to confirm the correlation between demographic factors and time use. We explored three k-modes clustering methods-backward, forward, and a new hybrid approach-to delineate the occupancy patterns based on demographics. Using the probability of cluster membership for each population segment and a time inhomogeneous Markov chain to generate activity transition probabilities for each cluster, we simulated 50,000 schedules per segment and validated them against the ATUS data. The hybrid method produced the most socio-demographically differentiated clusters while demonstrating comparable performance to other approaches, with an overall root mean square error of 0.12 for both weekday and weekend schedules. Thus, the hybrid method, where each cluster is dominated by certain demographic segments and occupancy patterns, offers more modeling versatility in terms of scenario analysis. The new workflow improves the socio demographic differentiation of energy consumption by considering differences in time use. This approach enables future research on demographically segmented time of use (TOU) energy consumption, including impacts of TOU utility bills and rate analysis, long-run marginal emissions, and energy retrofits.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic Hardware

Spiking Neural Networks (SNNs) are an emerging computation model that uses event-driven activation and bio-inspired learning algorithms. SNN-based machine learning programs are typically executed on tile-based neuromorphic hardware platforms, where each tile consists of a computation unit called a crossbar, which maps neurons and synapses of the program. However, synthesizing such programs on an off-the-shelf neuromorphic hardware is challenging. This is because of the inherent resource and latency limitations of the hardware, which impact both model performance, e.g., accuracy, and hardware performance, e.g., throughput. We propose DFSynthesizer, an end-to-end framework for synthesizing SNN-based machine learning programs to neuromorphic hardware. The proposed framework works in four steps. First, it analyzes a machine learning program and generates SNN workload using representative data. Second, it partitions the SNN workload and generates clusters that fit on crossbars of the target neuromorphic hardware. Third, it exploits the rich semantics of the Synchronous Dataflow Graph (SDFG) to represent a clustered SNN program, allowing for performance analysis in terms of key hardware constraints such as number of crossbars, dimension of each crossbar, buffer space on tiles, and tile communication bandwidth. Finally, it uses a novel scheduling algorithm to execute clusters on crossbars of the hardware, guaranteeing hardware performance. We evaluate DFSynthesizer with 10 commonly used machine learning programs. Our results demonstrate that DFSynthesizer provides a much tighter performance guarantee compared to current mapping approaches.

Computer Science↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Smart thermostat data-driven U.S. residential occupancy schedules and development of a U.S. residential occupancy schedule simulator

Occupancy schedule is one of the key inputs in Building Energy Modeling (BEM) to reflect the interaction between buildings and occupants. Over the past decades, standardized occupancy schedules, developed mainly by engineering rule-of-thumb, have been widely used in BEM due to its simplicity and lack of real measured occupancy data. However, the BEM community has recognized their association with uncertainty and reliability in simulation results from BEM. This study introduces representative occupancy schedules in the U.S. residential buildings, derived from a large smart thermostat dataset and time-series K-means clustering, and an open-source tool to generate a stochastic residential occupancy schedule. Over 90,000 residential occupancy schedules were estimated from the ecobee Donate Your Data dataset. Then, the representative occupancy schedules were identified through clustering. This study further investigated the impacts of three parameters (day, house type, and state) on residential occupancy schedules. Then, a tool, the Residential Occupancy Schedule Simulator (ROSS), is developed using the representative occupancy schedules derived in this study. Details of this tool are presented in this paper. In conclusion, the derived representative occupancy schedules and the ROSS tool can help improve the energy modeling of residential buildings.

42 ENGINEERING↗

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗

Evaluating HPC Scheduling Strategies for Urgent Workloads

Scientific computing centers increasingly face workloads with diverse urgency requirements, driven by applications that demand rapid or even immediate execution. Appropriately configured scheduling policies can significantly improve both user satisfaction and overall cluster utilization. In this work, we present a systematic analysis of scheduler configurations under scenarios where a fraction of jobs have urgent computing needs. We evaluate multiple job scheduling simulators, develop a lightweight job-submission emulation framework, and create tools to analyze and visualize the resulting scheduling data. Our study identifies key trade-offs between responsiveness, fairness, and efficiency, and offers a set of practical scheduling configurations (particularly for Slurm) that can be tailored to HPC environments supporting mixed-urgency workloads.

Maheshwari, Ketan [ORNL] (ORCID:000000033800662X)↗

Predicting runtime and resource utilization of jobs on integrated cloud and HPC systems

Recent advances in virtualization technologies used in cloud computing offer performance that closely approaches bare-metal levels. Combined with specialized instance types and high-speed networking services for cluster computing, cloud platforms have become a compelling option for high-performance computing (HPC). However, most current batch job schedulers in HPC systems are designed for homogeneous clusters and make decisions based on limited information about jobs and system status. Scientists typically submit computational jobs to these schedulers with a requested runtime that is often over- or under-estimated. More accurate runtime predictions can help schedulers make better decisions and reduce job turnaround times. Here, they can also support decisions about migrating jobs to the cloud to avoid long queue wait times in HPC systems.

97 MATHEMATICS AND COMPUTING↗

Oak Ridge Computing Academy: An HPC cluster deployment and management pilot

The High Performance Computing Technologies (HPCT) course is a hands-on High Performance Computing (HPC) cluster deployment and management training program offered as part of the International School for Advanced Studies (SISSA) and the International Center for Theoretical Physics (ICTP) Master in High Performance Computing (MHPC) specialization. Here, this training program introduces students to key concepts in cluster configuration. which include networking, software stack provisioning, job scheduling, and monitoring. The publicly available course materials feature several examples and underlying methods that are broadly applicable to cluster deployment and management. This paper discusses the design of a new workforce development program at the Oak Ridge National Laboratory that is based on HPCT, the Oak Ridge Computing Academy (ORCA). The ORCA pilot program was hosted by the Oak Ridge Leadership Computing Facility (OLCF) in Summer 2025. As a part of this discussion, HPCT and ORCA course contents and infrastructure are outlined, ORCA participant experiences are detailed, and potential opportunities for improvement are discussed.

Education↗

Preventive Power Outage Estimation Based on A Novel Scenario Clustering Strategy: Preprint

The increasing occurrence of extreme weather events is challenging the power grid operation. In front of the extreme weather, the system operator is responsible for estimating the power outage and scheduling the restoration resources. This paper proposes an outage evaluation framework to identify the possible unserved load profiles, vulnerable areas, and mobile energy adequacy. The predicted vulnerable lines of an outage prediction model tool are utilized to generate numerous faulted line scenarios. Next, each scenario's nodal unserved load profile is obtained by solving a three-phase restoration model that considers the schedule of repair crews and mobile energy resources. Then, a novel scenario clustering strategy is developed to cluster the unserved load profiles into multiple representative ones for straightforward analysis. Finally, case studies on a distribution system evaluate the damage level brought by extreme weather and verify the effectiveness of the proposed scenario clustering strategy.

mobile energy resources↗

A Hands-On Curriculum for Training in HPC Cluster Deployment and Management

This paper presents the design, methodology, and outcomes of the High-Performance Computing Technologies (HPCT) course, a hands-on training program focused on the system-side of HPC cluster deployment and administration. Delivered as part of the Master in High Performance Computing (MHPC) program, the course introduces students to key concepts in cluster configuration, including networking, software stack provisioning, job scheduling, and monitoring. Initially taught in person, the course was transitioned to an online format during the COVID-19 pandemic. This shift led to the development of openly available instructional material and a flipped-classroom approach that continues to support both in-person and hybrid delivery. All course materials are publicly available at www.hpc.temple.edu/mhpc/hpc-technology/index.html. By documenting the structure, infrastructure, and evolution of HPCT, this paper offers a model for accessible HPC system training that supports workforce development in computational science.

Posada Correa, Fernando [ORNL] (ORCID:000000022565↗

Shaping the FutureWorkforce: Challenges and Lessons Learned in HPC Education from National Labs and Computing Centers

Workforce training at national laboratories and computing centers is essential and typically falls into two categories: foundational training for newcomers and advanced training for experienced users. Foundational topics—such as version control, build systems, and basic HPC usage—are largely transferable across institutions, while cluster-specific training varies due to differences in hardware, job schedulers, and local workflows. Training on emerging technologies is split between hardware-specific content and broadly applicable programming paradigms. Here, to reduce redundancy and increase impact, national labs, computing centers, and vendors are collaborating through initiatives like the HPC Training Working Group to share best practices, co-develop materials, and broaden outreach. These coordinated efforts aim to make HPC training more accessible, scalable, and consistent across the community.

HPC↗

Informed Feature Selection for Data Clustering of CSP Plant Production

To make concentrating solar power (CSP) more cost competitive, rigourous optimizations must be run to improve plant design and operations. However, these optimizaitons rely on time consuming annual simulations that solve an electricity dispatch scheduling problem to maximize plant revenue. To reduce the runtime of annual dispatch simulations of CSP plants, a data clustering approach is utilized. This approach assumes that like days of revenue and electricity generation can be identified using weather and price data. Although weather and price are important factors for electricity production, this work investigates how thermal energy storage (TES) inventory at the beginning of a day, denoted as Si, can be used as a supplemental feature to group like days. A framework for creating and training a deep neural network to predict Si is proposed. This model is validated and assessed using eleven sets of testing data that were not used during training. Then, the data clustering approach is performed three seperate times with features of weather and price along with either Si from the neural network, Si from the full annual simulation, or no Si. Ultimately, the results suggest that using Si as an additional clustering feature improves the data clustering simulation accuracy by 1.4%.

Tuman, Matthew J. (ORCID:000900038772051X)↗