Search NASASearch

SEARCH · Search NASA

Results for “Scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak

Job Scheduler-Driven Power Gateway for High Performance Computing

Power gateways in the form of a microgrid can incorporate multiple distributed energy resources (DER) in either grid forming or grid following mode and support high performance computing (HPC) power profiles including the large load-follow requirements observed in multi-user HPC systems. The microgrid’s flexibility to operate in either grid forming or grid following mode and to actively switch between these modes enables baseline power from multiple non-baseline DER while maintaining high power quality metrics for the HPC system. But this enormous flexibility in demand response and time of use shifting is generally programmed independently of any integration with an HPC job scheduler which can better inform the load shaping by the microgrid. While there are many existing approaches where the HPC job scheduler takes in information from the grid to make queue scheduling decisions, this work takes the opposite view and explores a scheduler where the jobs in the queue can directly impact the settings of the grid. Several HPC scheduler strategies are tested where the jobs in the queue directly impact the settings of a microgrid designed for HPC operation which is driving a datacenter with three classes of HPC architectures. The scheduler operation is shown using a microgrid with 64 kW of solar capacity and 320 kWh of battery over a period of 21 days operating with significant low-follow swings, a throttled grid, cloudy conditions, switching between grid following and grid forming modes, and a wide range of battery states-of-charge all while maintaining high quality power metrics. The scheduler provides a mechanism for the job queue to directly impact a power gateway like a microgrid and to improve HPC power outcomes such as maximizing renewable energy usage

microgrid

HPC Digital Twins for Evaluating Scheduling Policies, Incentive Structures and their Impact on Power and Cooling

Schedulers are critical for optimal resource utilization in high-performance computing. Traditional methods to evaluate sched- ulers are limited to post-deployment analysis, or simulators, which do not model associated infrastructure. In this work, we present the first-of-its-kind integration of scheduling and digital twins in HPC. This enables what-if studies to understand the impact of parameter configurations and scheduling decisions on the physical assets, even before deployment, or regarching changes not easily realizable in production. We (1) provide the first digital twin framework extended with scheduling capabilities, (2) integrate various top-tier HPC systems given their publicly available datasets, (3) implement extensions to integrate external scheduling simulators. Finally, we show how to (4) implement and evaluate incentive structures, as- well-as (5) evaluate machine learning based scheduling, in such novel digital-twin based meta-framework to prototype scheduling. Our work enables what-if scenarios of HPC systems to evaluate sustainability, and the impact on the simulated system.

Maiterth, Matthias [ORNL] (ORCID:000000018698460X)

The relative influences of hydrologic information and dams’ hydropower scheduling decisions on electricity price forecasts

Price dynamics in wholesale electricity markets are driven by supply and demand. In markets with hydroelectric dams, the timing and amount of hydropower offered can influence prices in similar ways to wind and solar power. Unlike variable renewable energy, however, the supply of hydropower in wholesale markets is a function of both water availability and operational decisions at dams. Dam operators maximize revenues in wholesale markets by aligning generation with the periods of highest expected prices, and these scheduling decisions may in turn influence prices. Here, we examine the relative importance of two types of information in predicting forward electricity prices: a) water availability at dams, in the form of short-to-medium-range hydrological forecasts; and b) hourly scheduling decisions at dams. Using softly coupled hydrologic, hydropower scheduling, and power systems models spanning the U.S. Western Interconnection, we quantify the importance of hydrologic forecast accuracy in correctly predicting wholesale electricity prices and compare this with the influence of dam operators’ own hourly scheduling decisions on realized market prices. We find that aligning hydropower generation schedules with the periods of high forecasted prices causes larger, inadvertent price forecast errors than imperfect hydrologic forecasts. This suggests that knowledge of how water is managed by dam operators within the week is more important than weekly inflow forecast errors when predicting forward electricity prices. Our findings have implications for optimal hydropower scheduling by region. Specifically, accounting for price effects is critical in markets dominated by hydropower capacity.

Electricity markets

Link Scheduling in Satellite Networks via Machine Learning Over Riemannian Manifolds

Low Earth Orbit (LEO) satellites play a crucial role in enhancing global connectivity, serving a complementary solution to existing terrestrial systems. In wireless networks, scheduling is a vital process that allocates time-frequency resources to users for interference management. However, LEO satellite networks face significant challenges in scheduling their links towards ground users due to the satellites’ mobility and overlapping coverage. This paper addresses the dynamic link scheduling problem in LEO satellite networks by considering spatio-temporal correlations introduced by the satellites’ movements. The first step in the proposed solution involves modeling the network over Riemannian manifolds, thanks to their representation as symmetric positive definite matrices. We introduce two machine learning (ML)-based link scheduling techniques that model the dynamic evolution of satellite positions and link conditions over time and space. To accurately predict satellite link states, we present a recurrent neural network (RNN) over Riemannian manifolds, which captures spatio-temporal characteristics over time. Furthermore, we introduce a separate model, the convolutional neural network (CNN) over Riemannian manifolds, which captures geometric relationships between satellites and users by extracting spatial features from the network topology across all links. Simulation results demonstrate that both RNN and CNN over Riemannian manifolds deliver comparable performance to the fractional programming-based link scheduling (FPLinQ) benchmark. Remarkably, unlike other ML-based models that require extensive training data, both models only need 30 training samples to achieve over 99% of the sum rate while maintaining similar computational complexity relative to the benchmark.

42 ENGINEERING

Consumer safety-oriented scheduling of rotating power outages during heat waves

Extreme heat events have widespread effects on power systems, reducing available generation capacity, limiting transmission capabilities, and causing unusual demand patterns on the consumer side. As these combined effects expose bulk transmission systems to potential large-scale blackouts, utilities may be required to schedule and apply rotating outages, by temporarily and alternately disconnecting distribution substations to reduce overload. However, utilities lack mechanisms to inform these events, exacerbating the negative effects of heat waves on affected communities. This paper introduces a novel framework for scheduling rotating outages during heat waves while considering impacts on consumers’ safety. Instead of random sequential load shedding, we propose a methodology to rotate power outages considering a metric that quantifies the indoor overheating risk of groups of consumers during a power outage. The overheating risk is derived from a detailed building simulation using CityBES, where the buildings are modeled based on available data—use type, year built, floor area, number of stories, location—while presence of air conditioning and occupancy are calibrated from smart meter data. Based on the metric, an algorithm to schedule the rotating outages is applied to prioritize feeders for disconnection at each hour according to their overheating risk to meet a utility load reduction target. Applied to two substations and seven feeders in the Portland General Electric territory, the results show that this approach effectively leads to the lowest overheating risk during the resulting outage schedules, with an average 10.1% lower overheating compared to uninformed schedules.

Building thermal simulation

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]

Evaluating HPC Scheduling Strategies for Urgent Workloads

Scientific computing centers increasingly face workloads with diverse urgency requirements, driven by applications that demand rapid or even immediate execution. Appropriately configured scheduling policies can significantly improve both user satisfaction and overall cluster utilization. In this work, we present a systematic analysis of scheduler configurations under scenarios where a fraction of jobs have urgent computing needs. We evaluate multiple job scheduling simulators, develop a lightweight job-submission emulation framework, and create tools to analyze and visualize the resulting scheduling data. Our study identifies key trade-offs between responsiveness, fairness, and efficiency, and offers a set of practical scheduling configurations (particularly for Slurm) that can be tailored to HPC environments supporting mixed-urgency workloads.

Maheshwari, Ketan [ORNL] (ORCID:000000033800662X)

Joint scheduling of energy, fast and primary frequency response reserves in integrated transmission–distribution networks

Inverter-based distributed energy resources (DERs) connected to distribution networks (DNs) can provide fast frequency support, but their reserve deliverability depends on feeder constraints and differs from synchronous primary frequency response (PFR). Existing transmission–distribution coordination studies usually treat reserve generically or neglect feeder-level feasibility, while frequency-security scheduling studies rarely represent distribution feeders explicitly. This paper develops a bi-level day-ahead scheduling framework for integrated transmission–distribution networks that jointly clears energy, transmission-side PFR, and distribution-side fast frequency response (FFR) under exogenous hourly inertia and largest-loss inputs from an external unit commitment (UC) schedule. The transmission problem is modeled with DC-optimal power flow (OPF) and closed-form second-order cone (SOC) frequency-security constraints, whereas each DN is represented by a reserve-aware branch-flow AC-OPF so that scheduled fast reserves remain deliverable during activation. The bi-level problem is reformulated through Karush–Kuhn–Tucker (KKT) conditions into a mixed-integer SOC program, and a penalty term is used to tighten the distribution-network relaxation. In the reduced test system, lower exogenous inertia increased the required primary response from 179.64 MW to 191.08 MW, distribution-side fast response reduced total frequency-response procurement by up to 4.9%, and neglecting distribution constraints overstated the combined distribution-side energy and reserve award by up to 18%. In the expanded study, the largest case was solved in 2.02 s with a 0.00% optimality gap. Time-domain simulations kept the frequency nadir above 59.0 Hz in all tested hours. These results demonstrate the value of fast-response modeling and distribution-feasible reserve delivery in coordinated market clearing.

Noh, Seung-Gil

Robustness: The Missing Ingredient in Generation Scheduling

This article highlights robustness as an essential factor to cope with the ever-increasing levels of uncertainty in generation scheduling under significant renewable energy penetration, as is the case in Brazil and Spain. To that end, robust generation scheduling is framed within the different optimization-based approaches that are available for uncertainty handling. In addition, the suitability of robust optimization to accommodate practical security criteria in generation scheduling is also emphasized. Interestingly, this article points out the existence of an effective algorithm allowing the discovery of critical or so-called umbrella scenarios, which paves the way for the implementation of robust generation scheduling in industry practice.

Street, Alexandre

Correlating and Simulating Socio-Demographically Driven Residential End-Use Activity Schedules

Incorporating socio-demographic and behavioral considerations into decision-support tools is crucial for identifying gaps and addressing consumer needs to ensure reliable and affordable energy solutions. In energy simulation models, the correlation between socio-demographics and time-use behavior is not well-captured. Thus, we developed a large-scale simulation workflow to generate schedules for 10 residential activities across 24 population segments defined by age, income, and employment status. Using pre-pandemic 2015-2019 American Time Use Survey (ATUS) data, we used ANOVA to confirm the correlation between demographic factors and time use. We explored three k-modes clustering methods-backward, forward, and a new hybrid approach-to delineate the occupancy patterns based on demographics. Using the probability of cluster membership for each population segment and a time inhomogeneous Markov chain to generate activity transition probabilities for each cluster, we simulated 50,000 schedules per segment and validated them against the ATUS data. The hybrid method produced the most socio-demographically differentiated clusters while demonstrating comparable performance to other approaches, with an overall root mean square error of 0.12 for both weekday and weekend schedules. Thus, the hybrid method, where each cluster is dominated by certain demographic segments and occupancy patterns, offers more modeling versatility in terms of scenario analysis. The new workflow improves the socio demographic differentiation of energy consumption by considering differences in time use. This approach enables future research on demographically segmented time of use (TOU) energy consumption, including impacts of TOU utility bills and rate analysis, long-run marginal emissions, and energy retrofits.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Energy Scheduling-based Operating Envelopes including a Distribution System Branch Screening Algorithm

This paper presents an energy scheduling-based formulation for computing operating envelopes including a distribution branch screening algorithm, termed DBS-ES. The contribution of the paper is two-fold: firstly, it presents an innovative methodology for calculating operating envelopes using energy scheduling (baseline), and secondly, it enhances this methodology by incorporating a custom distribution branch screening algorithm (DBS-ES). The custom algorithm leverages power system knowledge to reduce both model build time and total processing time while maintaining the same scheduling results as the baseline. The effectiveness of the proposed approach is demonstrated through experiments on the IEEE13, IEEE123, and EPRI Secondary test feeders. Results highlight a 24.5% decrease in model build time and an 8.17% decrease in total processing time when using DBS-ES compared to the baseline, specifically for the IEEE123 test feeder. Additionally, the paper briefly discusses the influence of utility-controlled storage on computing operating envelopes, noting a general incre

24 POWER TRANSMISSION AND DISTRIBUTION

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC

Adaptive PID Gain Scheduling Control for Hydropower Turbine Using Neural CDE and Stochastic Distribution Shaping

This paper introduces a gain-scheduling PID controller design strategy for hydroturbine frequency control mode. This scheme first uses real data to learn the nonlinear dynamics of the hydroturbine using neural controlled differential equations and then perturbs the obtained nonlinear system at different equilibrium points, based on which a static output feedback adaptive dynamic programming algorithm is then used to optimize the PID gains for each equilibrium point. Moreover, a continuous-time version of stochastic distribution control is proposed to further fine-tune the optimized PID gains. Finally, the controller is obtained by implementing linear interpolation between the optimized PID control gains. The simulation results show that the proposed gain-scheduling PID controller can control a larger range of operation points compared with the given fixed PID controller and the baseline method. Compared with the given fixed PID controller, the proposed gain-scheduling PID controller can regulate hydroturbine frequency against disturbances induced by power-load variation with over 50% less overshoot for some operation points.

13 HYDRO ENERGY

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau

Flexible Resource Scheduler for FAST-DERMS (FRS-FASTDERMS) v0.9

The Flexible Resource Scheduler is a hierarchical controller that manages the distributed energy resources in a distribution substation or distribution feeder to provide a firm commitment of power flow at the substation or feeder head to be scheduled in transmission-level markets as an aggregated demand resource. It is the reference controller for the FAST-DERMS Architecture, developed in tandem with the architecture under the DOE FAST-DERMS project. It is comprised of a day-ahead stochastic optimization, which schedules substation power flow and reserves, a intra-hour MPC, which generates dispatch base points for DER, and a real-time PID controller maintaining that dispatches DER to maintain the substation power around the base points. The repository also includes a representative aggregator controller, and all of the necessary components to run a simulation using PNNL's GridAPPS-D software with the controller.

MacDonald, Jason [Lawrence Berkeley National Labor

Scheduler Modeling of Distributed Energy Resources for Providing Ancillary Services

Distribution energy resources (DERs) have been integral components of modern power systems, and their capability to provide grid services has been widely studied. To promote the deployment of these resources in providing grid services in real-world utility operations, this paper proposes a day-ahead scheduler model for a distribution system connected DER plant. A certain amount of generation capacity of this DER plant is reserved for frequency services, and some ancillary services for the distribution system-including peak load reduction, voltage regulation, and power factor control-are integrated into the model. The model is tested on a real-world distribution system. From the simulation results, the energy and reserve schedule of the solar photovoltaic (PV) unit and battery energy storage system (BESS) can be determined, and voltage and power factor are well maintained. Additionally, in order to demonstrate the specific characteristics of the co-located and hybrid operation modes for the PV and BESS, these two modes are analyzed both theoretically and through real-time simulation. Simulation results show that most of PV's variability is transferred to the net power in the co-located mode, whereas it is transferred to the BESS in the hybrid mode. This proposed scheduler model and the comparison of co-located and hybrid modes can provide practical guidance for the applications of DER plant in the real-world utility.

14 SOLAR ENERGY