Search NASA⌕ Search

SEARCH · Search NASA

Results for “Workflow management systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

AI-assisted detector design for the EIC (AID(2)E)

Artificial Intelligence is poised to transform the design of complex, large-scale detectors like ePIC at the future Electron Ion Collider. Featuring a central detector with additional detecting systems in the far forward and far backward regions, the ePIC experiment incorporates numerous design parameters and objectives, including performance, physics reach, and cost, constrained by mechanical and geometric limits. This project aims to develop a scalable, distributed AI-assisted detector design for the EIC (AID(2)E), employing state-of-the-art multiobjective optimization to tackle complex designs. Supported by the ePIC software stack and using G EANT 4 simulations, our approach benefits from transparent parameterization and advanced AI features. The workflow leverages the PanDA and iDDS systems, used in major experiments such as ATLAS at CERN LHC, the Rubin Observatory, and sPHENIX at RHIC, to manage the compute intensive demands of ePIC detector simulations. Tailored enhancements to the PanDA system focus on usability, scalability, automation, and monitoring. Ultimately, this project aims to establish a robust design capability, apply a distributed AI-assisted workflow to the ePIC detector, and extend its applications to the design of the second detector (Detector-2) in the EIC, as well as to calibration and alignment tasks. Additionally, we are developing advanced data science tools to efficiently navigate the complex, multidimensional trade-offs identified through this optimization process.

97 MATHEMATICS AND COMPUTING↗

Accelerating Resilience of the Community through Holistic Engagement and use of Renewables (ARCHER) Planning Framework

The primary objective of the Accelerating Resilience of the Community through Holistic Engagement and Use of Renewables (ARCHER) initiative was to identify and incorporate the unique variations in energy burden, social vulnerability, living conditions, and access to essential services that differ across communities. By accounting for these localized factors—down to the neighborhood level—the project supports more targeted and effective investments in community resilience. The framework seeks to establish practical planning guidance, methods, and performance measures for community energy resilience, integrate community-level and electric utility system resilience planning, and assess its effectiveness through comparison with conventional and operational planning approaches. A key component of the project was its data exchange platform, which is used to evaluate and demonstrate the tools, methodologies, and planning approaches developed through ARCHER. This open-source platform enables developers and vendors of distribution and outage management systems to build upon the research by incorporating its concepts into their own tools and workflows. This capability is enabled by the transparent availability of data, functional requirements, and the underlying information model. The project yielded several important insights. First, meaningful engagement with communities is essential to achieving comprehensive resilience outcomes. Second, resilience planning is most effective when electric grid considerations and broader community needs are addressed in a coordinated manner. Third, the use of platforms that allow for real-time input from communities can enhance utility responsiveness during restoration activities. Fourth, a structured and systematic planning approach can successfully translate ARCHER concepts into practice. Fifth, the development of an integrated metric that reflects both grid performance and community impacts provides a more holistic basis for evaluating resilience. Finally, incorporating community engagement and equity considerations into grid operations is critical, particularly during severe weather events that result in extended outages.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Evolution of the ATLAS TDAQ online software framework towards Phase-II upgrade: Use of Kubernetes as an orchestrator of the ATLAS Event Filter computing farm

The ATLAS experiment at the LHC at CERN continuously evolves its TDAQ system to meet the challenges of new physics goals and technological advancements. As ATLAS prepares for the Phase-II Run 4 of the LHC, significant enhancements in the TDAQ Controls and Configuration (TDAQ-CC) tools have been designed to ensure efficient data collection, processing, and management. This abstract presents the evolution of ATLAS TDAQ-CC system leading up to Phase-II Run 4. As part of the evolution towards Phase-II, Kubernetes has been chosen to orchestrate the Event Filter (EF) farm. By leveraging Kubernetes, ATLAS can dynamically allocate computing resources, scale processing capacity in response to changing data taking conditions and ensure high availability of data processing services. The integration of the Kubernetes with the TDAQ Run Control framework enables perfect synchronisation between the experiment’s data acquisition components and the computing infrastructure. We will discuss the architectural considerations and implementation challenges involved in Kubernetes integration with the ATLAS TDAQ-CC system. We will highlight the benefits of using Kubernetes as an EF farm orchestrator, including improved resource utilization, enhanced fault tolerance, and simplified deployment and management of data processing workflows. In addition, we will report on the extensive testing of Kubernetes that was conducted using a farm of 2500 servers within the experiment data taking environment, demonstrating its scalability and robustness in handling the demands of the ATLAS TDAQ system for Phase-II. The adoption of Kubernetes represents a significant step forward in the evolution of ATLAS TDAQ-CC system, aligning with industry best practices in container orchestration.

Corso Radu, Alina [Univ. of California, Irvine, CA↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

A Microservices Architecture Toolkit for Interconnected Science Ecosystems

Microservices architecture is a promising approach for developing reusable scientific workflow capabilities for inte- grating diverse resources, such as experimental and observational instruments and advanced computational and data management systems, across many distributed organizations and facilities. In this paper, we describe how the INTERSECT Open Architec- ture leverages federated systems of microservices to construct interconnected science ecosystems, review how the INTERSECT software development kit eases microservice capability develop- ment, and demonstrate the use of such capabilities for deploying an example multi-facility INTERSECT ecosystem.

Brim, Michael↗

Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

Device-to-device variability in experimental noise critically impacts reproducibility, especially in automated, high-throughput systems like additive manufacturing farms. While manageable in small labs, such variability can escalate into serious risks at larger scales, such as architectural 3D printing, where noise may cause structural or economic failures. This contribution presents a noise-aware decision-making algorithm that quantifies and models device-specific noise profiles to manage variability adaptively. It uses distributional analysis and pairwise divergence metrics with clustering to choose between single-device and robust multi-device Bayesian optimization strategies. Unlike conventional methods that assume homogeneous devices or enforce generic robustness, the proposed framework explicitly determines whether shared optimization across devices is appropriate based on the degree of inter-device noise heterogeneity. This enables improved performance, reproducibility, and efficiency. An experimental case study involving three nominally identical 3D printers (same brand, model, and close serial numbers) demonstrates reduced redundancy, lower resource usage, and improved reliability, along with improved convergence stability and solution quality through the selection of the appropriate optimization strategy based on the degree of inter-device noise heterogeneity. Overall, this framework establishes a general approach for precision- and resource-aware optimization in scalable, automated experimental platforms, demonstrated here on a representative multi-device 3D printing case study.

Schenk, Christina↗

Integrating the PanDA Workload Management System with the Vera C. Rubin Observatory

The Vera C. Rubin Observatory will produce an unprecedented astronomical data set for studies of the deep and dynamic universe. Its Legacy Survey of Space and Time (LSST) will image the entire southern sky every three to four days and produce tens of petabytes of raw image data and associated calibration data over the course of the experiment’s run. More than 20 terabytes of data must be stored every night, and annual campaigns to reprocess the entire dataset since the beginning of the survey will be conducted over ten years. The Production and Distributed Analysis (PanDA) system was evaluated by the Rubin Observatory Data Management team and selected to serve the Observatory’s needs due to its demonstrated scalability and flexibility over the years, for its Directed Acyclic Graph (DAG) support, its support for multi-site processing, and its highly scalable complex workflows via the intelligent Data Delivery Service (iDDS). PanDA is also being evaluated for prompt processing where data must be processed within 60 seconds after image capture. This paper will briefly describe the Rubin Data Management system and its Data Facilities (DFs). Finally, it will describe in depth the work performed in order to integrate the PanDA system with the Rubin Observatory to be able to run the Rubin Science Pipelines using PanDA.

79 ASTRONOMY AND ASTROPHYSICS↗

Automating Testing of DUNE Electronics via a Finite State Machine

The Deep Underground Neutrino Experiment (DUNE) is a flagship international collaboration designed to study neutrinos tiny, nearly massless particles that may hold answers to fundamental questions about the Universe. Fermilab s Robotic Test Stand (RTS) plays a critical role in ensuring the quality of approximately 50,000 Application-Specific Integrated Circuit (ASIC) chips that will be used in DUNE s massive liquid argon detectors. These electronics will be inside the cryostat; therefore, they will need to have a high yield of working chips and low noise. To improve the automation and reliability of the RTS, this project focused on designing and implementing a Python-based finite state machine (FSM) to manage chip handling workflows. The FSM was developed as a modular software framework to coordinate robotic arm movements, manage chip tray positions, and monitor system states during testing. Key features include robust error handling routines, a pause/resume system for safe mid-cycle interruptions, and a simulation mode for iterative testing without hardware dependencies. The system was designed to prepare for seamless integration with RTS hardware components such as the robotic arm and vision system. This integration will streamline collaboration and enable efficient deployment of updates across the six total institutions performing testing. The outcomes of this internship contribute to Fermilab s mission to advance high-energy physics and support the DOE s national goals by directly improving the testing of equipment to be used in DUNE. The project also provided valuable experience in software design and contributing to the success of DUNE.

Kang, Caleb [William Rainey Harper Coll.]↗

Automating Testing of DUNE Electronics via a Finite State Machine

The Deep Underground Neutrino Experiment (DUNE) is a flagship international collaboration designed to study neutrinos—tiny, nearly massless particles that may hold answers to fundamental questions about the Universe. Fermilab’s Robotic Test Stand (RTS) plays a critical role in ensuring the quality of approximately 50,000 Application-Specific Integrated Circuit (ASIC) chips that will be used in DUNE’s massive liquid argon detectors. These electronics will be inside the cryostat; therefore, they will need to have a high yield of working chips and low noise. To improve the automation and reliability of the RTS, this project focused on designing and implementing a Python-based finite state machine (FSM) to manage chip handling workflows. The FSM was developed as a modular software framework to coordinate robotic arm movements, manage chip tray positions, and monitor system states during testing. Key features include robust error handling routines, a pause/resume system for safe mid-cycle interruptions, and a simulation mode for iterative testing without hardware dependencies. The system was designed to prepare for seamless integration with RTS hardware components such as the robotic arm and vision system. This integration will streamline collaboration and enable efficient deployment of updates across the six institutions performing testing. The outcomes of this internship contribute to Fermilab’s mission to advance high-energy physics and support the DOE’s national goals by directly improving the testing of equipment to be used in DUNE. The project also provided valuable experience in software design and contributing to the success of DUNE.

Kang, Caleb [Fermilab]↗

Frontier Job-Centric Telemetry Dataset

Comprehensive analysis of high-performance computing (HPC) systems requires linking workload execution to system behavior. This kind of analysis is vital for diagnosing performance issues, managing capacity, detecting anomalous workloads, and understanding how applications interact with system hardware. This job-centric telemetry dataset unifies scheduler job records with node-level measurements, enabling direct association between workloads and their corresponding power, thermal, and performance characteristics. It contains sanitized, scheduler related metadata for 152,400 individual jobs that ran on the Frontier supercomputer and ended on selected days throughout 2024 and 2025, a subpopulation of ~6.8% of the total number of allocated jobs with non-zero run time on the system over that same period. Each is linked with files that contain telemetry time series records of the power utilization and temperature behavior of its allocated nodes and their processors during the run time of the job. Where available, a portion of the job files also contain network performance time series. Jobs are sampled from select days that reflect normal levels of user activity and possess job size distributions with large numbers of leadership class jobs (>20% of Frontier nodes). Jobs in this dataset attempt to best represent successful user workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

AI-Ready Control System for the Fermilab Accelerator Complex

Reliable, high-intensity operation of the Fermilab Accelerator Complex is critical to the success of the Long-Baseline Neutrino Facility and Deep Underground Neutrino Experiment. We describe the requirements and infrastructure necessary to support routine use of artificial intelligence and machine learning (AI/ML) in the accelerator control system. Three capabilities are identified: a machine learning operations (MLOps) framework standardizing the lifecycle of AI/ML automation from data management through deployment and monitoring; a data quality framework defining and enforcing standards required to build trustworthy AI/ML applications; and workflow integration with large language models to assist physicists, engineers, and operators with information retrieval, code development, and routine analysis. Use cases spanning beam diagnostics, beam control, and support system automation illustrate the technical requirements across the complex.

43 PARTICLE ACCELERATORS↗

A galactic approach to neutron scattering science

Neutron scattering science is leading to significant advances in our understanding of materials and will be key to solving many of the challenges that society is facing today. Improvements in scientific instruments are actually making it more difficult to analyze and interpret the results of experiments due to the vast increases in the volume and complexity of data being produced and the associated computational requirements for processing that data. New approaches to enable scientists to leverage computational resources are required, and Oak Ridge National Laboratory (ORNL) has been at the forefront of developing these technologies. We recently completed the design and initial implementation of a neutrons data interpretation platform that allows seamless access to the computational resources provided by ORNL. For the first time, we have demonstrated that this platform can be used for advanced data analysis of correlated quantum materials by utilizing the world's most powerful computer system, Frontier. In particular, we have shown the end-to-end execution of the DCA++ code to determine the dynamic magnetic spin susceptibility χ(q, ω) for a single-band Hubbard model with Coulomb repulsion U/t = 8 in units of the nearest-neighbor hopping amplitude t and an electron density of n = 0.65. The following work describes the architecture, design, and implementation of the platform and how we constructed a correlated quantum materials analysis workflow to demonstrate the viability of this system to produce scientific results.

97 MATHEMATICS AND COMPUTING↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

GDSA framework, a computational framework for complex modeling problems in radioactive waste management

This paper details a computational framework to produce automated, graphical workflows, and how this framework can be deployed to support complex modeling problems like those in nuclear engineering. Key benefits of the framework include: automating previously manual workflows; intuitive construction and communication of workflows through a graphical interface; and automated file transfer and handling for workflows deployed across heterogeneous computing resources. This paper demonstrates the framework's application to probabilistic post-closure performance assessment of systems for deep geologic disposal of nuclear waste. However, the framework is a general capability that can help users running a variety of computational studies.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

FY 2025 Multidimensional Data Correlation Platform: Unified Software Architecture for Advanced Materials and Manufacturing Technologies Data Management and Processing

The Advanced Materials and Manufacturing Technologies (AMMT) program continues to advance a data-driven approach to demonstrate the utility of additive manufacturing for fabricating components for nuclear applications. A key scientific goal is to leverage data to better understand manufacturing outcomes and thereby improve the performance, reliability, and lifespan of nuclear components. Ultimately, this effort supports the development of standards for certification and qualification of additively manufactured components, enabling broader industry adoption. In support of this objective, the AMMT program is building and deploying a data management platform to record, index, analyze, and make available the manufacturing data generated across the AMMT program. In FY 2023, the team conceptualized the architecture of the platform and, in FY 2024, deployed the first functional version at the Oak Ridge National Laboratory (ORNL) Manufacturing Demonstration Facility (MDF). In FY 2025, the platform was officially opened to all AMMT members. To enable this expansion, core modifications and enhancements were developed, including improvements to the user interface and workflows for data entry and retrieval. Most notably, robust security and access control mechanisms were implemented to protect data and manage information sharing. This effort featured a logging system, protected views, and controlled access mechanisms. This report documents these enhancements and the transition of the platform into program-wide use.

36 MATERIALS SCIENCE↗

MSD CoP Webinar: "Advances in MSD-LIVE to Support the MSD Community of Practice"

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Advances in MSD-LIVE to Support the MSD Community of Practice Presenters: Casey Burleyson and Zoe Guillen (Pacific Northwest National Laboratory) Abstract: The MultiSector Dynamics Living, Intuitive, Value-adding, Environment (MSD-LIVE; msdlive.org) is a cloud-based data management system and advanced computing platform that enables MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and workflows within the MSD Community of Practice. Recently, several high-profile datasets have attracted many new users to MSD-LIVE. This webinar has two goals: 1) To refamiliarize the MSD community and new users with the components of the platform (e.g., the data repository, model training notebooks, and data dashboards) and to highlight examples of how these components are advancing MSD science and 2) To demonstrate new features in v3 of the platform, released in late 2025. The main new feature in v3 is the ability to interactively explore data in MSD-LIVE without downloading it. MSD-LIVE users can now click a button in our data repository and launch a blank Jupyter notebook with access to the underlying data on AWS. Users can use the notebook to write analysis, visualization, or subsetting routines that process the data directly on the AWS cloud. We also added a GitHub integration feature that allows users to share analysis or visualization code they develop with the community of MSD-LIVE users. The webinar will wrap up with a look at what's coming next for MSD-LIVE in 2026. Moderator: Patrick M. Reed (MSD CoP Facilitation Team) This webinar was held on: May 12th, 2026 from 1-2 PM EST.

Open Science↗

INTEGRATION OF DATA ANALYTICS WITH SYSTEM HEALTH PROGRAMS

Industry equipment reliability and asset management programs are essential elements that help ensure the safe and economical operation of nuclear power plants. The effectiveness of these programs is addressed in several industry developed and regulatory programs. However, these programs have proven to be labor intensive and expensive. There is an opportunity to significantly enhance the collection, analysis, and use of this information to provide more cost-effective plant operation. Additionally, there is an acute industry need to leverage advanced technology to reduce costs and improve operational effectiveness. The goal of this paper is to provide effective and efficient analytical methods and tools to support risk-informed decisions for the equipment reliability and asset management programs at nuclear power plants. This is accomplished by creating a direct bridge between component health/lifecycle data and decision making (e.g., maintenance scheduling and project prioritization). Here we are supporting typical system engineer decisions regarding maintenance activity scheduling and component ageing management. This is performed in a risk-informed context where herein the term “risk” is broadly constructed to include both plant reliability and economics. This framework combines data analytics tools to analyze equipment reliability data with risk-informed methods designed to support system engineer decisions (e.g., maintenance and replacement schedules, optimal maintenance posture) in a customizable workflow. A challenge is that the structure of this workflow strongly depends on the decision that needs to be made, the type of data available, and the constraints that need to be considered. Current methods are designed to provide specific answers to specific problems; however, these methods might prove to be inadequate even when problem settings slightly change (e.g., different types of requirements, additional dependencies between system reliability and economics). We tackled this challenge by designing framework in a flexible and modular fashion such that the user can assemble and customize his/her own workflow that integrates SSC economic lifecycle models (e.g., maintenance and replacement costs), system reliability models, and optimization methods.

97 - MATHEMATICS AND COMPUTING↗

Active learning path-dependent properties using a cloud-based materials acceleration platform

Solid state materials are central to many modern technologies in which a given material may be exposed to a variety of environments. The material properties often vary with the sequence of environments in an irreversible manner, resulting in a quintessential path-dependency in experimental observables. While sequential learning techniques have been effectively deployed for accelerating learning of state properties of materials, they often use a consistent environment path in all experiments. To elevate such techniques for making optimal decisions in experimental investigations of path-dependent properties, we introduce an iterated expected information gain acquisition function that optimizes over entire experimental trajectories. This approach is implemented within a cloud-based Materials Acceleration Platform architecture utilizing an event-driven stateful broker coupled with remote HELAO (Hierarchical Experimental Laboratory Automation and Orchestration) instances and an AI science manager. The platform's efficacy was demonstrated through a case study optimizing multi-step spectro-electrochemical experiments to identify optically stable potential windows in (Co–Ni–Sb)O z metal oxides. The system successfully integrated AI-driven experiment design, remote laboratory automation, and cloud-based data infrastructure, validating the platform's capability for managing complex, adaptive, path-dependent workflows in materials discovery.

Guevarra, Dan [California Institute of Technology ↗