Search NASASearch

DOE OSTI · 3017617

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

Abstract

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Guan, Wen [Brookhaven National Laboratory (BNL), Upton, NY (United States)] (ORCID:0000000255485194), Maeno, Tadashi [Brookhaven National Laboratory (BNL), Upton, NY (United States)] (ORCID:0000000309011817), Alekseev, Aleksandr [Univ. of Texas, Arlington, TX (United States)] (ORCID:000000017025432X), Megino, Fernando Harald Barreiro [Univ. of Texas, Arlington, TX (United States)] (ORCID:0000000219990222), De, Kaushik [Univ. of Texas, Arlington, TX (United States)] (ORCID:0000000256474489), Karavakis, Edward [Brookhaven National Laboratory (BNL), Upton, NY (United States)] (ORCID:0000000257295167), Klimentov, Alexei [Brookhaven National Laboratory (BNL), Upton, NY (United States)] (ORCID:0000000327484829), Korchuganova, Tatiana [Univ. of Pittsburgh, PA (United States)] (ORCID:0000000157928182), Lin, FaHui [Univ. of Texas, Arlington, TX (United States)] (ORCID:0009000868207696), Nilsson, Paul [Brookhaven National Laboratory (BNL), Upton, NY (United States)] (ORCID:0000000268487463), Wenaus, Torre [Brookhaven National Laboratory (BNL), Upton, NY (United States)] (ORCID:000000028678893X), Yang, Zhaoyu [Brookhaven National Laboratory (BNL), Upton, NY (United States)] (ORCID:0009000987612547), Zhao, Xin [Brookhaven National Laboratory (BNL), Upton, NY (United States)] (ORCID:0000000321750452). 2026-01-24. iDDS: intelligent distributed dispatch and scheduling for workflow orchestration. https://doi.org/10.1140/epjc%2Fs10052-025-15275-7

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

TANTE: Time-adaptive operator learning via neural Taylor expansion

Operator learning for time-dependent partial differential equations (PDEs) has seen rapid progress in recent years, enabling efficient approximation of complex spatiotemporal dynamics. However, most existing methods rely on fixed time step sizes during rollout, which limits their ability to adapt to varying temporal complexity and often leads to error accumulation. In this work, we propose the Time-Adaptive Transformer with Neural Taylor Expansion (TANTE), a novel operator-learning framework that produces continuous-time predictions with adaptive step sizes. TANTE predicts future states by performing a Taylor expansion at the current state, where neural networks learn both the higher-order temporal derivatives and the local radius of convergence. This allows the model to dynamically adjust its rollout based on the local behavior of the solution, thereby reducing cumulative error and improving computational efficiency. We demonstrate the effectiveness of TANTE across a wide range of PDE benchmarks, achieving superior accuracy and adaptability compared to fixed-step baselines, delivering accuracy gains of 60-80 % and speed-ups of 30-40 % at inference time.

97 MATHEMATICS AND COMPUTING

Structured illumination for surface-resolved grazing-incidence X-ray scattering

Grazing-incidence (GI) scattering techniques are widely used to characterize thin films, offering high surface sensitivity and insight into morphology and structure. However, these approaches typically provide statistical averaged information due to elongated footprint or limited spatial resolution due to beam size. Here we introduce a method that combines structured illumination with GI X-ray scattering and leverages our computational imaging approach to resolve local structural details. We demonstrate that our method captures local features of an organic semiconductor thin film without the need for sample rotation as in tomography. The method expands GI techniques from statistical averaging to high-resolution imaging, thereby providing the capability for detailed analysis of local material properties, such as domain shape, orientation and polymorphism, which are critical for advancing material design towards more efficient and tailored materials.

97 MATHEMATICS AND COMPUTING