Search NASA⌕ Search

Engineering topics

Malviya, Addi Thakur

Publications and source records attributed to Malviya, Addi Thakur.

Towards a Software Development Framework for Interconnected Science Ecosystems

The innovative science of the future must be multi-domain and interconnected to usher in the next generation of “self-driving” laboratories enabling consequential discoveries and transformative inventions. Such a disparate and interconnected ecosystem of scientific instruments will need to evolve using a system-of-systems (SoS) approach. The key to enabling application integration with such an SoS will be the use of Software Development Kits (SDKs). Currently, SDKs facilitate scientific research breakthroughs via algorithmic automation, databases and storage, optimization and structure, pervasive environmental monitoring, among others. However, existing SDKs lack instrument-interoperability and reusability capabilities, do not effectively work in an open federated architectural environment, and are largely isolated within silos of the respective scientific disciplines. Inspired by the scalable SoS framework, this work proposes the development of INTERSECT-SDK to provide a coherent environment for multi-domain scientific applications to benefit from the open federated architecture in an interconnected ecosystem of instruments. This approach will decompose functionality into loosely coupled software services for interoperability among several solutions that do not scale beyond a single domain and/or application. Furthermore, the proposed environment will allow operational and managerial inter-dependence while providing opportunities for the researchers to reuse software components from other domains and build universal solution libraries. We demonstrate this research for microscopy use-case, where we show how INTERSECT-SDK is developing the tools necessary to enable advanced scanning methods and accelerate scientific discovery.

Malviya, Addi Thakur↗

Domain-Specific Type-Safe APIs for Hierarchical Scientific Data with Modern C++

General-purpose library application programming interfaces (APIs) for self-describing hierarchical scientific data storage, such as the HDF5 and NetCDF libraries, are traditionally of runtime nature. Runtime errors for entry existence and data types are typically caught later in the development process of higher-level application-specific APIs. In this paper, we propose exploiting modern C++ metaprogramming features to add compile-time type-safety to improve the interaction with a well-defined metadata-rich scientific schema in domain-specific hierarchical datasets. We tackle two aspects of common use: (i) direct data access, (ii) flexible “in-memory” index models for efficient search and data processing. The proposed APIs use C++17’s template type auto deduction features, C++11’s enum class for type-safety and C-style preprocessor macros for generative templated code. We showcase the pros and cons of our initial work on the standard NeXus schema used for annotating and storing experimental neutron scattering data at several facilities around the world on top of HDF5. Extendable compile-time type-safe APIs are a desirable feature that could be indexed by any modern integrated development environment (IDE). Hence, such APIs can help ease the learning curve for domain scientists using a less error-prone software interaction to enhance the findability of their data without resorting to a domain-specific language (DSL).

Godoy, William↗

Real-World Experiences Adopting Workflows at Exascale on the ExaAM Project

The purpose of this study is to discuss the experiential lessons associated with adopting scientific workflows in the Exascale Additive Manufacturing project (ExaAM) through the lens of Perceived Characteristic of Innovation (PCI). Besides the implementation, the factors we considered critical to the adoption of the workflow are provenance, sustainable automation, implementation challenges, and integration/compatibility challenges. Through conversations and interviews among the program managers, project leads, and software engineers, we have developed critical insight and strategies to overcome the obstacles and augment the successful adoption and long-term use of these workflows in ExaAM and beyond. We hope our work will pave the way for others in the research community to develop and use workflows in their respective science domains.

Malviya, Addi Thakur↗

A Survey on Sustainable Software Ecosystems to Support Experimental and Observational Science at Oak Ridge National Laboratory

In the search for a sustainable approach for software ecosystems that supports experimental and observational science (EOS) across Oak Ridge National Laboratory (ORNL), we conducted a survey to understand the current and future landscape of EOS software and data. This paper describes the survey design we used to identify significant areas of interest, gaps, and potential opportunities, followed by a discussion on the obtained responses. The survey formulates questions about project demographics, technical approach, and skills required for the present and the next five years. The study was conducted among 38 ORNL participants between June and July of 2021 and followed the required guidelines for human subjects training. We plan to use the collected information to help guide a vision for sustainable, community-based, and reusable scientific software ecosystems that need to adapt effectively to: i) the evolving landscape of heterogeneous hardware in the next generation of instruments and computing (e.g. edge, distributed, accelerators), and ii) data management requirements for data-driven science using artificial intelligence.

Bernholdt, David↗

Use of Event-Time Embeddings via RNN to Discern Novel Event Sequences in EHRs

In highly configurable health information technology (HIT) systems, such as VistA of the Veterans Health Administration, the variations in how the system is used among different healthcare facilities and how the data are recorded can be significant. Despite the successful standardization of care efforts, some of these variations can be indicative of HIT hazards and demand further investigation. In this work, we implemented a recurrent neural network (RNN) architecture to learn clinical provider order sequences and their temporal dynamics while predicting the orders' terminal state. We demonstrate model performance and provide a use case for the model discerning novel event sequences. This model is proposed to find novel event sequences in an operational environment.

Ozmen, Ozgur↗

Research Software Engineering Efforts for DataFlow: FY2021 Developments

DataFlow is a web application that helps scientific data to flow from one source location to another destination location. DataFlow helps scientists easily capture scientific metadata associated with an experiment and transmit both metadata and experimental data to a designated, centralized data storage resource. This report describes the software engineering efforts and architecture of the project for the fiscal year 2021 developments. We hope it effectively communicates findings from our work, challenges we have overcome, and how we will continue our development of DataFlow in the future.

42 ENGINEERING↗

Research Software Engineering Efforts for DataFed: FY2021 Developments

DataFed is a scientific data management system for big data providing simple and uniform data access, organization, discovery, and sharing within and across scientific facilities - with the goals of enhancing productivity and scientific reproducibility. We hope to aid in communicating development contributions to DataFed over the last year and any planning we can provide for the future developments in the next year

42 ENGINEERING↗

Comparative Assessment of Data-driven Process Models in Health Information Technology

Process mining for conformance analysis consists of comparing a reference process model against a data-driven process model generated via log files from information technology systems. However, in the absence of a complete reference process model, we found no suggested approaches in the literature to address the need for evaluating process conformance among different healthcare facilities to assess standardization of care. Our goal is to find similarities and dissimilarities in data-driven process models among US Veterans Health Administration (VHA) facilities that can be indicative of patient safety issues. Our hypothesis was that the analysis would not produce statistically significant differences in outcome. We present a unique implementation of conformance analysis in process mining that consists of combining process mining, process mapping and statistical metrics. We illustrate our approach by applying it to the analysis of two clinical radiology order process models generated from healthcare data provided by two similar facilities in the VHA. The comparative assessment showed that about 70% of the orders completed successfully and 30% were not completed due to policy and duplications. Our analysis found a good statistical correlation between both facilities, as the Spearman’s correlation coefficient between facilities for the frequency of cases per total hours was 0.87879, for the frequency of cases by state transition was 0.79702 and for the throughput time per state transition was 0.63582. Additional statistical analyses using the Mann-Whitney U test and the root mean square error both produced values that were not significant. The foregoing approach validated our hypothesis by demonstrating a good statistical correlation of data describing the flow of clinical radiology orders absent a credible reference model. Finding good agreement between both facilities was important in confirming that the clinical orders flow in a similar manner, suggesting standardization of care.

97 MATHEMATICS AND COMPUTING↗