Search NASA⌕ Search

Engineering topics

Wolf, Matthew

Publications and source records attributed to Wolf, Matthew.

Towards a Software Development Framework for Interconnected Science Ecosystems

The innovative science of the future must be multi-domain and interconnected to usher in the next generation of “self-driving” laboratories enabling consequential discoveries and transformative inventions. Such a disparate and interconnected ecosystem of scientific instruments will need to evolve using a system-of-systems (SoS) approach. The key to enabling application integration with such an SoS will be the use of Software Development Kits (SDKs). Currently, SDKs facilitate scientific research breakthroughs via algorithmic automation, databases and storage, optimization and structure, pervasive environmental monitoring, among others. However, existing SDKs lack instrument-interoperability and reusability capabilities, do not effectively work in an open federated architectural environment, and are largely isolated within silos of the respective scientific disciplines. Inspired by the scalable SoS framework, this work proposes the development of INTERSECT-SDK to provide a coherent environment for multi-domain scientific applications to benefit from the open federated architecture in an interconnected ecosystem of instruments. This approach will decompose functionality into loosely coupled software services for interoperability among several solutions that do not scale beyond a single domain and/or application. Furthermore, the proposed environment will allow operational and managerial inter-dependence while providing opportunities for the researchers to reuse software components from other domains and build universal solution libraries. We demonstrate this research for microscopy use-case, where we show how INTERSECT-SDK is developing the tools necessary to enable advanced scanning methods and accelerate scientific discovery.

Malviya, Addi Thakur↗

Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path Forward

The ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows.

Mehta, Kshitij↗

F*** workflows: when parts of FAIR are missing

The FAIR principles for scientific data (Findable, Accessible, Interoperable, Reusable) are also relevant to other digital objects such as research software and scientific workflows that operate on scientific data. The FAIR principles can be applied to the data being handled by a scientific workflow as well as the processes, software, and other infrastructure which are necessary to specify and execute a workflow. The FAIR principles were designed as guidelines, rather than rules, that would allow for differences in standards for different communities and for different degrees of compliance. There are many practical considerations which impact the level of FAIR-ness that can actually be achieved, including policies, traditions, and technologies. Because of these considerations, obstacles are often encountered during the workflow lifecycle that trace directly to shortcomings in the implementation of the FAIR principles. Here, we detail some cases, without naming names, in which data and workflows were Findable but otherwise lacking in areas commonly needed and expected by modern FAIR methods, tools, and users. We describe how some of these problems, all of which were overcome successfully, have motivated us to push on systems and approaches for fully FAIR workflows.

Wilkinson, Sean↗