Search NASA⌕ Search

SEARCH · Search NASA

Results for “Use Cases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

ASEAN Technical Exchange Workshop for System Operators, Regulators, and Policymakers

This presentation provides an in-depth exploration of power system planning, cross-border electricity trading, and battery energy storage systems (BESS), offering actionable insights for system operators, regulators, and policymakers. The first section delves into power system planning and analysis, focusing on capacity expansion models and resource adequacy studies, including their role in optimizing system efficiency, managing emissions, and addressing system reliability risks. Key considerations, such as integration of transmission into generation planning and the forecasting versus optimization of customer distributed energy resources (DER) technologies, are explored. The session highlights critical trade-offs in spatial granularity and model runtimes, as well as the feasibility of aligning distribution investments with capacity expansion efforts. The second section examines cross-border electricity trading, with an emphasis on resource adequacy concepts such as reliability targets, loss of load expectation (LOLE), and planning reserve margins (PRM). Case studies on reserve market design and coordination across US regions provide insights into improving reserve deliverability and managing interregional power balance and congestion. This section also addresses market-to-market congestion management, including advanced strategies for high-voltage direct current (HVDC) optimization and ancillary service delivery. Finally, the presentation covers the rapid evolution of Battery Energy Storage Systems (BESS), highlighting their operational growth, regulatory frameworks, and use cases in grid flexibility, energy storage, and reliability. The discussion focuses on the benefits of BESS for system stability, resilience, and integration of renewable energy, offering insights into its role as a vital component in the transition toward a more sustainable and flexible grid. Key performance parameters, such as throughput, round-trip efficiency, and state of charge, are also examined.

25 ENERGY STORAGE↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

Explicit block encodings of boundary value problems for many-body elliptic operators

Simulation of physical systems is one of the most promising use cases of future digital quantum computers. In this work we systematically analyze the quantum circuit complexities of block encoding the discretized elliptic operators that arise extensively in numerical simulations for partial differential equations, including high-dimensional instances for many-body simulations. When restricted to rectangular domains with separable boundary conditions, we provide explicit circuits to block encode the many-body Laplacian with separable periodic, Dirichlet, Neumann, and Robin boundary conditions, using standard discretization techniques from low-order finite difference methods. To obtain high-precision, we introduce a scheme based on periodic extensions to solve Dirichlet and Neumann boundary value problems using a high-order finite difference method, with only a constant increase in total circuit depth and subnormalization factor. We then present a scheme to implement block encodings of differential operators acting on more arbitrary domains, inspired by Cartesian immersed boundary methods. We then block encode the many-body convective operator, which describes interacting particles experiencing a force generated by a pair-wise potential given as an inverse power law of the interparticle distance. This work provides concrete recipes that are readily translated into quantum circuits, with depth logarithmic in the total Hilbert space dimension, that block encode operators arising broadly in applications involving the quantum simulation of quantum and classical many-body mechanics.

Kharazi, Tyler [University of California, Berkeley↗

An exploration of online-simulation-driven portfolio scheduling in Workflow Management Systems

Workflow Management Systems used to automate the execution of scientific workflow applications on parallel and distributed computing platforms must make scheduling decisions at runtime. A large number of workflow scheduling algorithms have been proposed in the literature, but often these algorithms are evaluated based on simplifying assumptions that may not hold in practice. Furthermore, published algorithm evaluation and/or comparison results are necessarily only for a subset of all possible scenarios, and thus may not include scenarios relevant to particular use-cases. Consequently, it is difficult for Workflow Management Systems (WMSs) developers to decide which scheduling algorithm should be implemented. To obviate this difficulty, one possible approach is to implement a portfolio of scheduling algorithms and select the most effective algorithm at runtime. One method for performing this selection is to run an online simulation for each algorithm in the portfolio. The algorithm that leads to the best performance, in simulation, is selected for future use. The above simulation-driven portfolio scheduling (SDPS) approach has been proposed in a few parallel and distributed computing contexts. The main objective of this work is to evaluate the feasibility and potential merit of SDPS if implemented in WMSs. Here we perform this evaluation using simulated WMS executions, where the simulations are instantiated from real-world platform and workflow configurations. Our main finding is that SDPS is on par with or outperforms an approach in which a single algorithm is used, where this algorithm is the one that performs best on average across all our experimental scenarios. Furthermore, we find that SDPS remains an attractive proposition even in the presence of high levels of simulation error and for simulators with relatively low levels of sophistication. In many of our experimental scenarios we find that mitigating simulation error at runtime can further improve performance. Finally, we show that simulation overhead can be made sufficiently low for SDPS to be feasible in practice.

97 MATHEMATICS AND COMPUTING↗

Numerical studies of collinear laser-assisted injection from a foil for plasma wakefield accelerators

We present a laser-assisted electron injection scheme for beam-driven plasma wakefield acceleration. The laser is collinear with the driver and triggers the injection of hot electrons into the plasma wake by interaction with a thin solid target. We present a baseline case using the AWAKE Run 2 parameters and then perform variations on key parameters to explore the scheme. It is found that the trapped witness electron charge may be tuned by altering laser parameters, with a strong dependence on the phase of the wake upon injection. Normalized emittance settles at the order of micrometres and varies with witness charge. The scheme is robust to misalignment, with a 1/10th plasma skin-depth offset ( 20 μ m for the AWAKE case) having a negligible effect on the final beam. The final beam quality is better than similar existing schemes, and several avenues for further optimization are indicated. The constraints on the AWAKE experiment are very specific, but the general principles of this mechanism can be applied to future beam-driven plasma wakefield accelerator experiments. Published by the American Physical Society 2024

Physics↗

Versatile TRISO fuel particle modeling in Bison

Tri-structural isotropic (TRISO) fuel particles are a key component in several previous and current reactors as well as in a variety of novel nuclear reactor designs. Interest in TRISO fuel is on the rise, necessitating considerable computer modeling of TRISO fuel behavior in order to support related design and licensing activities. The Bison nuclear fuel performance code, which offers a full set of capabilities for modeling TRISO fuels, makes it easier to explore the various important aspects of TRISO fuel behavior. One key advantage of Bison is its ability to create meshes in 1D, 2D, and 3D. Users can customize these meshes for specific geometries, mesh densities, and use cases. This enables a wide variety of analyses, including thermal, structural, mass diffusion, homogenization, and statistical failure analyses. Furthermore, the meshing capability simplifies analysts’ workflows. The inherent mesh generation capability eliminates the need for separate mesh-generating software and mesh file management. Also, the fact that the meshes are customizable makes it straightforward to automate an investigation over a range of geometric parameters or mesh densities. Here, the present paper highlights the ease with which Bison may be used to create meshes for both simple and relatively complex TRISO fuel particles, and it explores the types of analyses enabled by these meshes.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Air quality impacts from the development of unconventional oil and gas well pads: Air toxics and other volatile organic compounds

Unconventional oil and natural gas development (UOGD) has expanded rapidly across the United States in recent decades and raised concerns about associated air quality impacts. While significant effort has been made to quantify methane emissions, relatively few observations have been made of Volatile Organic Compounds (VOCs), especially during drilling and completion of new wells. Extensive air monitoring during development of several large, multi-well pads in Broomfield, Colorado, in the Denver-Julesburg Basin, provides a novel opportunity to examine changes in local air toxics and other VOC concentrations during well drilling and completions and production. These operations offer an especially useful case to study as several management practices were implemented to reduce emissions (e.g., electrified, grid-powered drill rigs and closed loop fluid handling systems to reduce truck traffic and limit fluid handling on the pad). With simultaneous measurements of methane and 50 VOCs from October 2018 to December 2022 at as many as 19 sites near well pads, in adjacent neighborhoods, and at a more distant reference location, we identify impacts from each phase of well development and production. Use of weekly, time-integrated canisters, a Proton Transfer Reaction Mass Spectrometer (PTR-MS), continuous photoionization detectors (PID) to trigger canister collection upon detection of VOC-rich plumes, and an instrumented vehicle, provided a powerful suite of measurements to characterize both transient plumes and longer-term changes in air quality. Prior to the start of well development, VOC gradients were small across Broomfield. Once drilling commenced, concentrations of oil and gas (O&G) related VOCs, including alkanes and aromatics, increased around active well pads. Concentration increases were clearly apparent during certain operations, including drilling, coil tubing/millout operations, and production tubing installation. Emissions of C 8 –C 10 n-alkanes during drilling operations highlighted the importance of VOC emissions from synthetic drilling mud chosen to reduce odor impacts. More than 90 samples were collected of transient plumes. Using composition measurements, meteorological data, and information about well pad activities, these plumes were connected with specific UOGD operations including drilling, flowback, and production equipment maintenance. The chemical signatures of these plumes differed by operation type (e.g., C 8 –C 10 n-alkanes constituted a larger fraction of measured VOCs in drilling-related plumes). Concentrations of individual, oil and gas-related VOCs in these plumes were often several orders of magnitude higher than in background air, with maximum ethane and benzene concentrations of 79,600 and 819 ppbv, respectively. Because these plumes typically impact a monitoring site for just several minutes, they are easily missed by slower-responding instruments. Study measurements highlight future emission mitigation opportunities during UOGD operations, including better control of emissions from shakers that separate drill cuttings from drilling mud, production separator maintenance operations, and periodic emptying of sand cans during flowback operations.

54 ENVIRONMENTAL SCIENCES↗

Improving photovoltaic hosting capacity of distribution networks with coordinated inverter control: A case study of the EPRI J1 feeder

Abstract Adding photovoltaic (PV) systems in distribution networks, while desirable for reducing the carbon footprint, can lead to voltage violations under high solar‐low load conditions. The inability of traditional volt‐VAr control in eliminating all the violations is also well‐known. This article presents a novel coordinated inverter control methodology that leverages system‐wide situational awareness to significantly improve hosting capacity (HC). The methodology employs a real‐time voltage‐reactive power (VQ) sensitivity matrix in an iterative linear optimizer to calculate the minimum reactive power intervention from PV inverters needed for mitigating over‐voltage without resorting to active power curtailing or requiring step voltage regulator setting changes. The algorithm is validated using the EPRI J1 feeder under an extensive set of realistic use cases and is shown to provide 3x improvement in HC under all scenarios.

Dalal, Dhaval [School of Electrical, Computer, and↗

Fully Homomorphic Encryption

This code implements a Fully Homomorphic Encryption (FHE) system, enabling secure computation on encrypted data without requiring decryption. It supports encryption, decryption, and homomorphic operations like matrix multiplication and addition. This code is adaptable for integrating FHE into linear-time invariant (LTI) systems, including digital control and filtering. With proper configuration from subject matter expertise, encrypted system parameters and signals can be manipulated to perform tasks like state updates, output calculations, and convolution in the encrypted domain. By preserving the structure of LTI systems while ensuring privacy, the framework facilitates secure applications in areas such as autonomous systems, signal processing, and industrial automation. The code initializes the encryption system using parameters provided in the env dictionary. These parameters include the ciphertext modulus, key dimension, plaintext fixed-point scaling factor, and noise bound. During initialization, a secret key is generated, which is essential for encrypting and decrypting data securely. The modular design allows users to tailor these parameters to specific use cases or security requirements. The code implements multiple cryptographic schemes. The learning with errors (LWE) encryption method encodes cleartext message to their plaintext fixed-point representation then encrypted into ciphertext space with additive noise. This noise ensures the security of the scheme, relying on the computational hardness of the LWE problem. The code also includes the Gentry-Sahai-Waters (GSW) scheme based off the LWE problem. Homomorphic matrix multiplication is performed between the LWE and GSW to encrypted data. This is achieved using a decomposition function on the LWE ciphertext during the multiplication operation. For higher-dimensional data, the code includes a method to encrypt entire matrices (GSWMat) using GSW encryption. These encrypted matrices can then be used for homomorphic matrix multiplications (MatMult). The decryption function uses the secret key to recover the original plaintext, removing the added noise and scaling that was originally applied during encryption.

Lois, Roberts [Idaho National Laboratory (INL), Id↗

Quantum Information Encoding and Decoding for Quantum Sensi

This two-year theory project focused on theoretical investigations of novel paradigms for quantum sensing, building on information encoding and techniques from quantum error correction, quantum computing and other quantum information domains. The outcomes facilitate quantum information technology development, especially at the interface of quantum computing and quantum sensing. The results of the project show new use cases and new paradigms for quantum sensing beyond what has so far been considered. One outcome shows how quantum sensing opens new opportunities for fundamental physics such as the capability of single graviton detection. Another outcome reveals a new application of NISQ quantum computers with error correction for metrology, building on recent advances in practical quantum error correction implementation. The third outcome of the project creates new paradigms of back-action-evading sensing inspired by collective quantum information encoding, which achieves quantum sensing beyond the quantum limit without the use of entanglement.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Proposed Classifications of Remote Operations for Nuclear Reactors Based on Physical and Cybersecurity Considerations

The incorporation of remote operations into reactor operations is a topic of high interest among advanced and small modular reactor (A/SMR) vendors, with some considering it essential to the success of their business models. However, remote operations are a concept novel to the nuclear industry. While various technical aspects of remote operations have been explored, a significant gap remains in understanding the security implications of integrating remote operations into reactor designs, particularly concerning the security requirements for remote-operations facilities and infrastructure. This report aims to address this gap by first defining classes of remote operation based on the extent of remote access to reactor control systems and grounded in the existing regulatory framework with compatible terminology. Secondly, the report outlines the physical and cybersecurity requirements applicable to remote-operations facilities and infrastructure at each defined class. These requirements are based on existing licensing frameworks provided by 10 Code of Federal Regulations (CFR) Part 50 and 10 CFR Part 52, as well as the upcoming A/SMR licensing framework in the proposed Part 53. The assessment focuses specifically on security regulations, such as 10 CFR Part 73, which includes provisions for both cybersecurity (§ 73.54) and physical security (§ 73.55). This report proposes five classes of remote reactor operations. Class 1 involves remote monitoring only, with no control over reactor systems. Class 2 allows for the remote issuance of allowlisted commands to the reactor facility. Class 3 extends control to non-safety-significant, non-safety-related, or not important to safety systems and equipment. Class 4 permits remote control of safety-significant systems. Finally, Class 5 allows remote control of safety-related systems. It is important to note that these classes were defined purely with functionality in mind, without considering the practicality or feasibility of implementation for each class under current or upcoming regulatory guidance. The intention behind this approach is to enable an assessment of which security requirements apply to each class, allowing readers to evaluate the implementation possibilities for their specific use cases. Following the definition of remote-operation classes, the report assesses the specific physical and cybersecurity requirements applicable to the remote-operations facility and infrastructure within each defined class. This includes defining the types and locations of operators that are possible at each class of operation and, based on operator type and location, as well as functionality within each class, outlining the physical and cybersecurity requirements. By detailing the security requirements by class, the report provides readers with the information needed to determine the type of security program they may need to implement for their desired concept of operation. The next contribution of this report was to assess the practicality of implementing each proposed class of remote operations based upon the security requirement assessment. In short, three of the five proposed remote-operation classes were found to possibly have a practical path forward to implementation under the U.S. regulatory framework. Class 1 remote operations are currently in use in the U.S. while Class 2 and 3 remote operations may be logistically possible to implement under the U.S. regulatory framework. The final two Classes, 4 and 5, would likely be logistically difficult, if not infeasible to implement within the current U.S. physical- and cybersecurity regulatory framework. Given the results of the feasibility assessment, an example architecture is proposed for both Class 2, remote allowlisted commands, and Class 3, remote control of non-safety systems as well as security implication assessments of each architecture. These example implementations are not meant to be prescriptive in terms of how Class 2 or Class 3 remote operations should be deployed; instead, they are intended to be informative to stakeholders on how Class 2 or Class 3 could potentially be applied in order to inform their system design. An example architecture for Class 1 remote monitoring was not provided as Class 1 in already in use in U.S. nuclear operations. Example architectures for Class 4 and Class 5 were not provided due to their assessment of being likely infeasible to implement. The final contribution is an assessment of the physical- and cybersecurity implications of introducing autonomous operations into an A/SMR. What was found was that the security implications can be separated into two cases. Autonomous operations supported by SSCs located only at the reactor site, and autonomous operations supported by SSCs outside of the reactor site. For the first case, the introduction of autonomous systems will likely not change the facility’s requirement to comply with existing cyber and physical security regulation

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Computational Workflows for Uncertainty-Quantified Nuclear Reactions: From Nuclear Theory Inputs to Astrophysical Reaction Rates

Reactions on unstable nuclei, particularly those on the neutron-rich side of stability, are important for both fundamental and applied physics. For fundamental science, the most prevalent use case is astrophysi cal nucleosynthesis by rapid neutron capture—the r-process—by which heavy nuclei are formed in extreme astrophysical environments, such as in supernovae and neutron star mergers; see, e.g., Refs. [1–3]. For ap plications, these processes are relevant for the interpretation of radiochemical data from historic nuclear tests, which contribute to our ability to certify the enduring stockpile in the absence of nuclear testing [4]; see Ref. [5] for a broader discussion of applications. However, reaction cross sections involving unsta ble species are generally poorly understood, for the simple reason that useful data become scarce as one moves away from stability. While there are avenues for improving the amount and quality of data for these species [6], one is fundamentally reliant on nuclear theory to make progress on these fields of study.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Wastewater reuse benefits for municipal complete retention lagoons: Life cycle assessment and dynamic modeling

Complete retention lagoons with wastewater reuse for agricultural purposes may offer sustainability advantages over alternative systems for small communities in semiarid regions. This study quantifies the environmental life cycle impact of adopting agriculture water reuse systems using case study data to estimate operating and building infrastructure impacts and spatial–temporal modeling to quantify resource trade-offs. Water reuse system benefits are highly dependent on supply–storage–demand dynamics. The relative size of irrigated agricultural land to the lagoon size was the most significant factor influencing site water application rates. The benefits are sensitive to changes in air emissions occurring from the agricultural land and further emphasize the importance of proper fertilizer management when adopting water reuse systems. Wastewater reuse from complete retention lagoons reduce life cycle GHG emissions, primarily through excavation reductions, offset fertilizer use, and especially from increased crop yields from wastewater reuse at previously rainfed sites.

54 ENVIRONMENTAL SCIENCES↗

Reduced-dimension Bayesian optimization for model calibration of transient vapor compression cycles

Development and calibration of first-principles dynamic models of vapor compression cycles (VCCs) is of critical importance for applications that include control design and fault detection and diagnostics. Nevertheless, the inherent complexity of models that are represented by large systems of differential–algebraic equations leads to significant challenges for model calibration processes that utilize classical gradient-based methods. Bayesian optimization (BO) is a sample-efficient and gradient-free approach using a probabilistic surrogate model and optimal search over a feasible parameter space. Despite the benefits of BO in reducing computational costs, challenges remain in dealing with a high-dimensional calibration task resulting from a large set of parameters that have significant impacts on system behavior and need to be calibrated simultaneously. This paper presents a reduced-dimension BO framework for calibrating transient VCCs models where the calibration space is projected to a low-dimensional subspace for accelerating convergence of the solution algorithm and consequently reducing the number of transient simulations. The proposed approach was demonstrated via two case studies associated with different VCC applications where 10 parameters were calibrated in each case using laboratory measurements. The reduced-dimension BO framework only required 1 / 8 th of the iterations associated with a standard BO method that deals with high-dimensional calibration parameters for converged solutions and yielded comparable accuracy. Furthermore, both calibrated models revealed significant accuracy improvements compared to uncalibrated models.

Ma, Jiacheng↗

Cyberattack Detection and Mitigation on Central Volt‐VAr Using Circuit Law and Machine Learning

ABSTRACT In a distribution grid, voltage is maintained within a nominal range through a Volt‐VAr function that controls capacitor banks, reactive power of distributed energy resources (DER), and on‐load tap changers (OLTC). Availability of communications helps with the implementation of central Volt‐VAr control; however, it also opens the system to cyberattacks, causing voltage disturbances. Previous work has shown the adverse impacts of false data injection (FDI) on the central Volt‐VAr control; however, very few works have studied methods to detect and mitigate FDI on Volt‐VAr control. This paper addresses gaps in the detection and mitigation of FDI on the measurement packets of a central Volt‐VAr control. This work uses a two‐stage algorithm for cyberattack detection since the accuracy of a single‐stage machine learning (ML)–based detection method decreases while dealing with unseen data. The first stage is based on the verification of measurements against circuit laws, and the second stage utilizes a tree search algorithm and an ML method to detect the falsified data. This paper compares long short‐term memory (LSTM) and bidirectional LSTM (BiLSTM) as the employed ML algorithms. Finally, the mitigation algorithm replaces the falsified data with the estimated output of the ML algorithm. The effectiveness of the proposed method is tested for several cases using the IEEE 13‐bus test system in PSCAD software.

Beikbabaei, Milad [Bradley Department of Electrica↗

ExaWorks software development kit: a robust and scalable collection of interoperable workflows technologies

Scientific discovery increasingly requires executing heterogeneous scientific workflows on high-performance computing (HPC) platforms. Heterogeneous workflows contain different types of tasks (e.g., simulation, analysis, and learning) that need to be mapped, scheduled, and launched on different computing. That requires a software stack that enables users to code their workflows and automate resource management and workflow execution. Currently, there are many workflow technologies with diverse levels of robustness and capabilities, and users face difficult choices of software that can effectively and efficiently support their use cases on HPC machines, especially when considering the latest exascale platforms. We contributed to addressing this issue by developing the ExaWorks Software Development Kit (SDK). The SDK is a curated collection of workflow technologies engineered following current best practices and specifically designed to work on HPC platforms. We present our experience with (1) curating those technologies, (2) integrating them to provide users with new capabilities, (3) developing a continuous integration platform to test the SDK on DOE HPC platforms, (4) designing a dashboard to publish the results of those tests, and (5) devising an innovative documentation platform to help users to use those technologies. Our experience details the requirements and the best practices needed to curate workflow technologies, and it also serves as a blueprint for the capabilities and services that DOE will have to offer to support a variety of scientific heterogeneous workflows on the newly available exascale HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Complete and Correct Transfer of Information (CACTI)

Many distributed systems, file transfer mechanisms, and message passing systems offer reliability mechanisms such as acknowledgements, retries, and durability. While these tools may be “good enough” for their typical use cases, they may not offer sufficient coverage for the wide range of faults that impact data transfers and communication. A gap in the reliability measures may lead to some small amount of data loss. Some high-consequence systems cannot tolerate the loss or corruption of even a single record. We present seven principles that will counter a wide range of faults and protect against data loss and corruption. These principles bring together lessons learned from a wide range of technologies and can inform appropriate system design and application usage. These principles will help readers reason on how prevent data loss in a multi-hop pipeline and how to properly use tools that may have a deficiency in reliability.

97 MATHEMATICS AND COMPUTING↗