Search NASA⌕ Search

Engineering topics

O’Leary, Patrick

Publications and source records attributed to O’Leary, Patrick.

The SENSEI Generic In Situ Interface: Tool and Processing Portability at Scale [Book Chapter]

One key challenge when doing in situ processing is the investment required to add code to numerical simulations needed to take advantage of in situ processing. Such instrumentation code is often specialized, and tailored to a specific in situ method or infrastructure. Then, if a simulation wants to use other in situ tools, each of which has its own bespoke API [4], then the simulation code team will quickly become overwhelmed with having a different set of instrumentation APIs, one per in situ tool or method. In an ideal situation, such instrumentation need happen only once, and then the instrumentation API provides access to a large diversity of tools. In this way, a data producer’s instrumentation need not be modified if the user desires to take advantage of a different set of in situ tools. The SENSEI generic in situ interface addresses this challenge, which means that SENSEI-instrumented codes enjoy the benefit of being able to use a diversity of tools at scale, tools that include Libsim, Catalyst, Ascent, as well as user-defined methods written in C++ or Python. SENSEI has been shown to scale to greater than 1M-way concurrency on HPC platforms, and provides support for a rich and diverse collection of common scientific data models. Furthermore, this chapter presents the key design challenges that enable tool and processing portability at scale, some performance analysis, and example science applications of the methods.

Bethel, E. Wes↗

Proximity Portability and in Transit , M-to-N Data Partitioning and Movement in SENSEI [Book Chapter]

In high-performance parallel in situ processing, the term in transit processing refers to those configurations where data must move from a producer to a consumer that runs on separate resources. In the context of parallel and distributed computing on an HPC platform one of the central challenges is to determine a mapping of data from producer ranks to consumer ranks. This problem is complicated by the heterogeneity that arises in producer-consumer pairs, such as when producer and consumer codes have different levels of concurrency, different scaling characteristics, or different data models. The resulting mapping and movement of data from M producer to N consumer ranks can have a significant impact on aggregate application performance, particularly when the data consumer requires only a subset of the overall data for its task. This chapter focuses on the design considerations that underlie SENSEI’s implementation to this challenging problem. These design considerations extend the core SENSEI architecture and include ideas like the need to accommodate flexibility in the choice of different partitioning methods, the ability for a data consumer to request and receive only the subset of data needed for its particular operation, and the ability to leverage any of several different data transport tools. The idea of proximity portability, being able to use different data transport methods as part of an in transit workflow, is illustrated through the use of three different transport layers where switching from one transport tool to another is accomplished with only a configuration file change. Here, the chapter also includes a performance analysis summary showing the performance gains that are possible in terms of multiple metrics, such as memory footprint, time to solution, and amount of data moved, when using optimized partitioners in an in transit setting, gains that are made possible by the implementation shaped by specific design considerations.

Bethel, E. Wes↗

Sandtank-ML: An Educational Tool at the Interface of Hydrology and Machine Learning

Hydrologists and water managers increasingly face challenges associated with extreme climatic events. At the same time, historic datasets for modeling contemporary and future hydrologic conditions are increasingly inadequate. Machine learning is one promising technological tool for navigating the challenges of understanding and managing contemporary hydrological systems. However, in addition to the technical challenges associated with effectively leveraging ML for understanding subsurface hydrological processes, practitioner skepticism and hesitancy surrounding ML presents a significant barrier to adoption of ML technologies among practitioners. In this paper, we discuss an educational application we have developed—Sandtank-ML—to be used as a training and educational tool aimed at building user confidence and supporting adoption of ML technologies among water managers. We argue that supporting the adoption of ML methods and technologies for subsurface hydrological investigations and management requires not only the development of robust technologic tools and approaches, but educational strategies and tools capable of building confidence among diverse users.

54 ENVIRONMENTAL SCIENCES↗

Catalyst Revised: Rethinking the ParaView in Situ Analysis and Visualization API

As in situ analysis goes mainstream, ease of development, deployment, and maintenance becomes essential, perhaps more so than raw capabilities. In this paper, we present the design and implementation of Catalyst, an API for in situ analysis using ParaView, which we refactored with these objectives in mind. Furthermore, our implementation combines design ideas from in situ frameworks and HPC tools like Ascent and MPICH.

97 MATHEMATICS AND COMPUTING↗