Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Enhanced Machine-Learning Flow for Microwave-Sensing Systems for Contaminant Detection in Food

The presence of foreign bodies in packaged food is a serious concern for both fnal consumers (allergies, injuries, choking) and food manufacturers (reputation and economic losses). In particular, low-density plastics, glass and wood splinters are hard to detect even by the most advanced X-ray imagers. One solution is Machine-Learning-based Microwave Sensing (MLMWS): a non-invasive, contactless, and real-time method which uses a machine-learning (ML) classifer to analyze the scattered microwaves from the irradiated target object. In this paper, we want to extend our previous work about contaminant detection in cocoa-hazelnut spread jars by proposing an enhanced ML flow to increase the accuracy of the ML classifier. For the first time in this case study, we use a multi-class classifier, we train it with scattering parameters measured at multiple microwave frequencies, with a new pre-processing scaler, data augmentation, quantization-aware training and a pruning schedule. The results show a contaminant detection multi-class accuracy of 94.167% with a latency of 26 µs when targeting an AMD/Xilinx Kria K26 FPGA. Finally, we released our datasets publicly to OpenML.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Magnetic Field Mapping Design for the MOLLER Spectrometer Magnets at Jefferson Lab

The Thomas Jefferson National Accelerator Facility (JLab) has developed a unique spectrometer system to study the weak interaction between electrons. The "Measurement of Lepton-Lepton Electroweak Reaction" (MOLLER) experiment, utilizing JLab's recent 12 GeV electron beam upgrade, is scheduled to operate for three years. Central to the MOLLER experiment are five water-cooled toroidal magnets, each with a unique geometry and seven-fold symmetry, designed to focus the particles. These magnets generate the magnetic field needed to separate incident beam electrons scattered from target electrons (Møller scattering) and protons (elastic e-p scattering) within a liquid hydrogen target. Here, this paper details the magnet field measuring technique developed to map all five MOLLER toroidal magnets at multiple locations inside and along the bore. It covers the design, mounting, and operation of the probe, along with the calibration procedure to determine the field and to prepare a field map for GEANT4 analysis. Additionally, the paper addresses the challenges of accurately measuring low magnetic fields.

Ghoshal, Probir K. [Thomas Jefferson National Acce↗

Computational Performance Bounds Prediction in Quantum Computing With Unstable Noise

Quantum computing has significantly advanced in recent years, boasting devices with hundreds of quantum bits (qubits), hinting at its potential quantum advantage over classical computing. Yet, noise in quantum devices poses significant barriers to realizing this supremacy. Understanding noise’s impact is crucial for reproducibility and application reuse; moreover, the next-generation quantum-centric supercomputing essentially requires efficient and accurate noise characterization to support system management (e.g., job scheduling), where ensuring correct functional performance (i.e., fidelity) of jobs on available quantum devices can even be higher-priority than traditional objectives. However, noise fluctuates over time, even on the same quantum device, which makes predicting the computational bounds for on-the-fly noise is vital. Noisy quantum simulation can offer insights but faces efficiency and scalability issues. Here, in this work, we propose a data-driven workflow, namely QuBound, to predict computational performance bounds. It decomposes historical performance traces to isolate noise sources and devises a novel encoder to embed circuit and noise information processed by a Long Short-Term Memory (LSTM) network. For evaluation, we compare QuBound with a state-of-the-art learning-based predictor, which only generates a single performance value instead of a bound. Experimental results show that the result of the existing approach falls outside of performance bounds, while all predictions from our QuBound with the assistance of performance decomposition better fit the bounds. Moreover, QuBound can efficiently produce practical bounds for various circuits with over 106 speedup over simulation; in addition, the range from QuBound is over 10× narrower than the state-of-the-art analytical approach.

Li, Jinyang [George Mason Univ., Fairfax, VA (Unit↗

Real-Time Gear-Shift Optimization for an Autonomous Wheel Loader

Off-road vehicles, such as wheel loaders, consume a significant amount of fuel during the transportation of materials. The gear-shifting process is crucial in fuel savings for transportation, and therefore, optimization of gearshifts is important for vehicle control. For an autonomous off-road vehicle, there is potential for more fuel savings by proper coordination of gearshift optimization and optimization of other control inputs. This brief proposes a new method for integrating gear-shifting into transportation optimization for an autonomous wheel loader to minimize fuel consumption. Furthermore, tests conducted on short loading cycles show that this method can save around 10%–20% of fuel on average compared with a conventional gearshift-scheduling method.

33 ADVANCED PROPULSION SYSTEMS↗

Safe Deep Reinforcement Learning for Robust Frequency and Voltage-Constrained Networked Microgrid Restoration

Here, this paper proposes a safe soft actor-critic reinforcement learning (RL) algorithm–based controller for networked microgrid restoration. It formulates the post black-start start as a finite-horizon constrained Markov decision process. The RL agent co-optimizes real and reactive power set-points for both grid-forming and grid-following inverters under explicit voltage and frequency constraints, while enforcing proper power sharing via the Mean Active Power Sharing Index (MPSI) and Mean Reactive Power Sharing Index (MQSI). Numerical results obtained on the IEEE 123-bus distribution system show that the proposed method achieves a mean voltage build-up time of 0.01 s without breaching the 5% sharing-violation budget under various load scenarios, considering MPSI and MQSI indices. These findings demonstrate that the proposed method yields fast and safe black-start schedules without resorting to heuristic penalties.

Selim, Alaa [Dartmouth College, Hanover, NH (Unite↗

Safe Deep Reinforcement Learning for Active Distribution System Model Predictive Control with EVs and DERs

The temporal and spatial mismatch between PV generation and electric vehicle (EV) charging and discharging may cause voltage violations in active distribution networks. Despite the widespread use of deep reinforcement learning (DRL) in power system optimization and control, it lacks guarantees on constraint satisfaction during both training and deployment. This paper proposes a Lagrangian-based safe DRL approach for model predictive control (MPC) of active distribution systems with large-scale integration of PVs, EVs, and energy storage systems (ESSs). A Transformer-LSTM time-series model is proposed to forecast EV charging demand, which is then formulated as a constraint to ensure charging requirements are met. Using this prediction, a Lagrangian-based safe soft actor-critic (SAC) framework is developed for real-time control in a three-phase unbalanced distribution system, enforcing voltage safety constraints while optimizing the cumulative net reward. By integrating the forecasting model with multi-period constraints, the proposed framework jointly coordinates PV systems, EV charging and discharging, and ESS scheduling within the MPC horizon. Numerical experiments on a modified IEEE 123-bus system with real-world data show that, under a high PV penetration scenario, the proposed method increases the net reward by 30.74% and reduces average voltage violations from 0.0011 p.u. to 0.0002 p.u. compared with standard SAC. Compared with the optimal power flow (OPF) approach, it achieves similar voltage security while yielding lower line losses. It also maintains real-time control capability, reducing operation latency to 53.21 ms per 15-minute control interval. The proposed method remains effective under varying PV/EV penetrations and load conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

X-Ray and Particle Detection With the Si(Li) Tracker Module of the GAPS Experiment

Here, this work describes the architecture and the experimental results from the characterization of the lithium-drifted silicon (Si(Li)) detector module, which constitutes the building block of the tracker in the general antiparticle spectrometer (GAPS) experiment to search for dark matter. The instrument is designed for the identification of low-energy cosmic anti-nuclei (antiprotons, antideuterons, and antihelium) to be performed during an Antarctic long-duration balloon flight scheduled for late 2025. The GAPS Si(Li) tracker, that is the core of the instrument, is the assembly of 252 modules, each comprised of four Si(Li) detectors and a full custom-integrated circuit designed for detector readout and produced in a commercial 180-nm planar CMOS technology. A general overview of the detector module architecture and its components is provided, together with a description of the test setup and the experimental results obtained from the characterization of the low-noise analog readout channel. In order to verify the effective operation of the entire module, results concerning the detection of X-rays from a 241Am source and cosmic muons are also provided.

Manghisoni, Massimo [Università di Bergamo (Italy)↗

Exploring Enhanced Dominant Resource Fairness Using Linear Programming Calculated Weights

Maintaining resource fairness while achieving optimization for various performance metrics such as resource utilization, turnaround time and job latency is a well-known resource scheduling challenge in cloud computing. Despite the significant progress made with the introduction of dominant resource fairness by Ghodsi et al., which ensures major allocation properties such as sharing incentive, strategy-proofness, envy-freeness and Pareto efficiency to be achieved

Yan, Bo [Binghamton University]↗

Detection of visible-wavelength aurora on Mars

Mars hosts various auroral processes despite the planet’s tenuous atmosphere and lack of a global magnetic field. To date, all aurora observations have been at ultraviolet wavelengths from orbit. We describe the discovery of green visible-wavelength aurora, originating from the atomic oxygen line at 557.7 nanometers, detected with the SuperCam and Mastcam-Z instruments on the Mars 2020 Perseverance rover. Near–real-time simulations of a Mars-directed coronal mass ejection (CME) provided sufficient lead-time to schedule an observation with the rover. The emission was observed 3 days after the CME eruption, suggesting that the aurora was induced by particles accelerated by the moving shock front. To our knowledge, detection of aurora from a planetary surface other than Earth has never been reported, nor has visible aurora been observed at Mars. This detection demonstrates that auroral forecasting at Mars is possible, and that during events with higher particle precipitation, or under less dusty atmospheric conditions, aurorae will be visible to future astronauts.

Science & Technology - Other Topics↗

Measurement of the inhomogeneity of the KATRIN tritium source electric potential by high-resolution spectroscopy of conversion electrons from $\mathbf {^{83m}}$Kr

Precision spectroscopy of the electron spectrum of the tritium β-decay near the kinematic endpoint is a direct method to determine the effective electron antineutrino mass. The KArlsruhe TRItium Neutrino (KATRIN) experiment aims to determine this quantity with a sensitivity of better than 0.3 eV (90% C.L.). An inhomogeneous electric potential in the tritium source of KATRIN can lead to distortions of the β-spectrum, which directly impact the neutrino-mass observable. This effect can be quantified through precision spectroscopy of the conversion-electrons of co-circulated metastable 83m Kr. Therefore, dedicated, several-weeks long measurement campaigns have been performed within the KATRIN data taking schedule. In this work, we infer the tritium source potential observables from these measurements, and present their implications for the neutrino-mass determination.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

RAP: Resource-aware Automated GPU Sharing for Multi-GPU Recommendation Model Training and Input Preprocessing

Ensuring high-quality recommendations for newly onboarded users requires the continuous retraining of Deep Learning Recommendation Models (DLRMs) with freshly generated data. To serve the online DLRM retraining, existing solutions use hundreds of CPU computing nodes designated for input preprocessing, causing significant power consumption that surpasses even the power usage of GPU trainers. To this end, we propose RAP, an end-to-end DLRM training framework that supports Resource-aware Automated GPU sharing for DLRM input Preprocessing and Training. The core idea of RAP is to accurately capture the remaining GPU computing resources during DLRM training for input preprocessing, achieving superior training efficiency without requiring additional resources. Specifically, RAP utilizes a co-running cost model to efficiently assess the costs of various input preprocessing operations, and it implements a resource-aware horizontal fusion technique that adaptively merges smaller kernels according to GPU availability, circumventing any interference with DLRM training. In addition, RAP leverages a heuristic searching algorithm that jointly optimizes both the input preprocessing graph mapping and the co-running schedule to maximize the end-to-end DLRM training throughput. The comprehensive evaluation shows that RAP achieves 78.3× speedup on average over CPU-based DLRM input preprocessing frameworks. In addition, the end-to-end training throughput of RAP is only 2.04% lower than the ideal case, which has no input preprocessing overhead.

Wang, Zheng↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

A Digital Twin of Scalable Quantum Clouds

Quantum computing has emerged as a transformative technology capable of solving complex problems beyond the limit of classical systems. The rapid development of quantum processors has led to the proliferation of cloud-based quantum computing services offered by platforms such as IBM, Google, and Amazon. These platforms introduce unique challenges in resource allocation, job scheduling, and multi-device orchestration as quantum workloads become increasingly complex. In this work, we present a digital twin of quantum cloud infrastructures: a framework designed to model and simulate the behavior of real quantum cloud systems. Developed in Python using the SimPy discrete-event simulation library, the framework replicates key aspects of quantum cloud environments, including detailed quantum device modeling, job lifecycle management, and job fidelity. It incorporates noise-aware fidelity estimation, making it the first of its kind to simulate superconducting gate-based quantum cloud systems at an administrative level with job fidelity. We present use cases as proof of concept, demonstrating that our quantum cloud simulation framework can act as a digital twin of a quantum cloud and support the modeling and implementation of practical systems.

Luo, Waylon [Kent State University]↗

CGSim: A Simulation Framework for Large Scale Distributed Computing Environment

Large-scale distributed computing infrastructures such as the Worldwide LHC Computing Grid (WLCG) require comprehensive simulation tools for evaluating performance, testing new algorithms, and optimizing resource allocation strategies. However, existing simulators suffer from limited scalability, hardwired algorithms, lack of real-time monitoring, and inability to generate datasets suitable for modern machine learning approaches. We present CGSim, a simulation framework for large-scale distributed computing environments that addresses these limitations. Built upon the validated SimGrid simulation framework, CGSim provides high-level abstractions for modeling heterogeneous grid environments while maintaining accuracy and scalability. Key features include a modular plugin mechanism for testing custom workflow scheduling and data movement policies, interactive real-time visualization dashboards, and automatic generation of event-level datasets suitable for AI-assisted performance modeling. We demonstrate CGSim’s capabilities through a comprehensive evaluation using production ATLAS PanDA workloads, showing significant calibration accuracy improvements across WLCG computing sites. Scalability experiments show near-linear scaling for multi-site simulations, with distributed workloads achieving 6 × better performance compared to single-site execution. The framework enables researchers to simulate WLCG-scale infrastructures with hundreds of sites and thousands of concurrent jobs within practical time budget constraints on commodity hardware.

Vatsavai, Sairam Sri [Brookhaven National Laborato↗

FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving

Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data parallelism (DP) maximizes throughput by running independent replicas, while tensor parallelism (TP) reduces per-request latency and pools memory for long-context inference. However, existing serving stacks typically commit to a static parallelism configuration at deployment; adapting to bursts, priorities, or long-context requests is often disruptive and slow. We present Flying Serving, a vLLM-based system that enables online DP-TP switching without restarting engine workers. Flying Serving makes reconfiguration practical by virtualizing the state that would otherwise force data movement: (i) a zero-copy Model Weights Manager that exposes TP shard views on demand, (ii) a KV Cache Adaptor that preserves request KV state across DP/TP layouts, (iii) an eagerly initialized Communicator Pool to amortize collective setup, and (iv) a deadlock-free scheduler that coordinates safe transitions under execution skew. Across three popular LLMs and realistic serving scenarios, Flying Serving improves performance by up to 4.79 × under high load and 3.47 × under low load while supporting latency- and memory-driven requests.

Gao, Shouwei [ORNL]↗

Pipeline Hydrogen Decarbonization and Repurposing Analyzers (P-HyDRAs)

The Pipeline Hydrogen Decarbonization and Repurposing Analyzers (P-HyDRAs) are a set of prototype computational tools for simulating and optimizing midstream natural gas pipeline system operations subject to location and time-dependent hydrogen blending. The models can accurately resolve dynamic gas flows through large-scale pipeline networks using non-ideal gas equations of state. The codes can be used as decision support for planning and design decisions involving intra-day energy flow schedules as well as spatiotemporal economic values of natural gas, hydrogen, and net energy delivered to consumers while ensuring that pipeline hydraulic limitations, gas compressor station constraints, operational factors, and pre-existing shipping contracts are satisfied. The inputs to the codes are a model of the pipeline system as well as time-series data that specify boundary conditions on the network. For optimization, the code module requires price and quantity offers for natural gas and hydrogen and price and quantity bids for energy, which are used as time-dependent constraints in an optimal control problem. The outputs are time-series data that provide a predictive simulation of gas flows, mass fractions, and pressures, or with additional degrees of freedom give an approximately optimal solution for gas injections/withdrawals, compressor settings, and sensitivities to the objective function that provide locational values of energy.

Zlotnik, Anatoly↗

HydroWIRES-PNNL/fisch

Forecast Informed Scheduler for Hydropower (FIScH)

Broman, Dan [Pacific Northwest National Laboratory↗

Generator Scorecard

Code built on Archive Walker. The Generator Scorecard analyzes power grid measurements to automate the process of evaluating the performance of generators. The tool includes three aspects: frequency response, voltage response, and voltage schedule tracking.

Follum, Jim↗