Search NASASearch

DOE OSTI · 3002321

Toward a persistent event-streaming system for high-performance computing applications

Dorier, Matthieu [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Gueroudji, Amal [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Hayot-Sasson, Valérie [Univ. of Chicago, IL (United States)]·Nguyen, Hai Duc [Argonne National Laboratory (ANL), Argonne, IL (United States); Univ. of Chicago, IL (United States)]·Ockerman, Seth [Univ. of Wisconsin, Madison, WI (United States)]·Souza, Renan Santos [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)]·Bicer, Tekin [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Pan, Haochen [Univ. of Chicago, IL (United States)]·Carns, Philip [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Chard, Kyle [Argonne National Laboratory (ANL), Argonne, IL (United States); Univ. of Chicago, IL (United States)]·Chard, Ryan [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Gonthier, Maxime [Argonne National Laboratory (ANL), Argonne, IL (United States); Univ. of Chicago, IL (United States)]·Huerta, Eliu [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Lenard, Ben [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Nicolae, Bogdan [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Patel, Parth [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Wozniak, Justin [Argonne National Laboratory (ANL), Argonne, IL (United States)]·Foster, Ian [Argonne National Laboratory (ANL), Argonne, IL (United States); Univ. of Chicago, IL (United States)]·Rao, Nageswara S. [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000234085941)·Ross, Robert B. [Argonne National Laboratory (ANL), Argonne, IL (United States)]

Abstract

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dorier, Matthieu [Argonne National Laboratory (ANL), Argonne, IL (United States)], Gueroudji, Amal [Argonne National Laboratory (ANL), Argonne, IL (United States)], Hayot-Sasson, Valérie [Univ. of Chicago, IL (United States)], Nguyen, Hai Duc [Argonne National Laboratory (ANL), Argonne, IL (United States); Univ. of Chicago, IL (United States)], Ockerman, Seth [Univ. of Wisconsin, Madison, WI (United States)], Souza, Renan Santos [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)], Bicer, Tekin [Argonne National Laboratory (ANL), Argonne, IL (United States)], Pan, Haochen [Univ. of Chicago, IL (United States)], Carns, Philip [Argonne National Laboratory (ANL), Argonne, IL (United States)], Chard, Kyle [Argonne National Laboratory (ANL), Argonne, IL (United States); Univ. of Chicago, IL (United States)], Chard, Ryan [Argonne National Laboratory (ANL), Argonne, IL (United States)], Gonthier, Maxime [Argonne National Laboratory (ANL), Argonne, IL (United States); Univ. of Chicago, IL (United States)], Huerta, Eliu [Argonne National Laboratory (ANL), Argonne, IL (United States)], Lenard, Ben [Argonne National Laboratory (ANL), Argonne, IL (United States)], Nicolae, Bogdan [Argonne National Laboratory (ANL), Argonne, IL (United States)], Patel, Parth [Argonne National Laboratory (ANL), Argonne, IL (United States)], Wozniak, Justin [Argonne National Laboratory (ANL), Argonne, IL (United States)], Foster, Ian [Argonne National Laboratory (ANL), Argonne, IL (United States); Univ. of Chicago, IL (United States)], Rao, Nageswara S. [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000234085941), Ross, Robert B. [Argonne National Laboratory (ANL), Argonne, IL (United States)]. 2025-09-17. Toward a persistent event-streaming system for high-performance computing applications. https://doi.org/10.3389/fhpcp.2025.1638203

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Shaping the FutureWorkforce: Challenges and Lessons Learned in HPC Education from National Labs and Computing Centers

Workforce training at national laboratories and computing centers is essential and typically falls into two categories: foundational training for newcomers and advanced training for experienced users. Foundational topics—such as version control, build systems, and basic HPC usage—are largely transferable across institutions, while cluster-specific training varies due to differences in hardware, job schedulers, and local workflows. Training on emerging technologies is split between hardware-specific content and broadly applicable programming paradigms. Here, to reduce redundancy and increase impact, national labs, computing centers, and vendors are collaborating through initiatives like the HPC Training Working Group to share best practices, co-develop materials, and broaden outreach. These coordinated efforts aim to make HPC training more accessible, scalable, and consistent across the community.

HPC

Understanding power and energy utilization in large scale production physics simulation codes

Power is an often-cited reason for the move to advanced architectures on the path to Exascale computing. Here, this is due to practical considerations related to delivering enough power to successfully site and operate these machines, as well as concerns about energy usage while running large simulations. Since obtaining accurate power measurements can be challenging, it may be tempting to use the processor thermal design power (TDP) as a surrogate due to its simplicity and availability. However, TDP is not indicative of typical power usage while running simulations. Using commodity and advanced technology systems at Lawrence Livermore and Sandia National Labs, we performed a series of experiments to measure power and energy usage in running simulation codes. These experiments indicate that large scale Lawrence Livermore simulation codes are significantly more efficient than a simple processor TDP model might suggest.

HPC

Dependable classical-quantum computing systems engineering

Increasing evidence suggests quantum computing (QC) complements traditional High-Performance Computing (HPC) by leveraging its unique capabilities, leading to the emergence of a new, hybrid paradigm, QHPC. However, this integration introduces new challenges, with dependability–defined by reproducibility, resiliency, and security and privacy–emerging as a central concern for building trustworthy systems that provide an advantage to the users. This paper proposes a framework for dependable QHPC system design, organized around these three pillars. We identify integration challenges, anticipate roadblocks, and highlight productive synergies across QC, HPC, cloud platforms, and network security. Drawing from both classical computing principles and quantum-specific insights, we present a roadmap for co-design that supports robust hybrid architectures. Our approach offers concrete metrics for assessing dependability, provides design guidance for engineers working at the QC-HPC interface, and surfaces new engineering questions around complexity, scale, and fault tolerance. Ultimately, designing for dependability is key to realizing practical, scalable QHPC systems and accelerating the broader quantum ecosystem capable of translating quantum promises into actual application delivery.

HPC