Search NASA⌕ Search

SEARCH · Search NASA

Results for “interactive HPC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Enabling Seamless Transitions from Experimental to Production HPC for Interactive Workflows

The evolving landscape of scientific computing requires seamless transitions from experimental to production HPC environments for interactive workflows. This paper presents a structured transition pathway developed at OLCF that bridges the gap between development testbeds and production systems. We address both technological and policy challenges, introducing frameworks for data streaming architectures, secure service interfaces, and adaptive resource scheduling for time-sensitive workloads and improved HPC interactivity. Our approach transforms traditional batch-oriented HPC into a more dynamic ecosystem capable of supporting modern scientific workflows that require near real-time data analysis, experimental steering, and cross-facility integration.

Etz, Brian [ORNL] (ORCID:0000000208554863)↗

Towards Interactive, Reproducible Analytics at Scale on HPC Systems

The growth in scientific data volumes has resulted in a need to scale up processing and analysis pipelines using High Performance Computing (HPC) systems. These workflows need interactive, reproducible analytics at scale. The Jupyter platform provides core capabilities for interactivity but was not designed for HPC systems. In this paper, we outline our efforts that bring together core technologies based on the Jupyter Platform to create interactive, reproducible analytics at scale on HPC systems. Our work is grounded in a real world science use case-applying geophysical simulations and inversions for imaging the subsurface. Our core platform addresses three key areas of the scientific analysis workflow-reproducibility, scalability, and interactivity. We describe our implemention of a system, using Binder, Science Capsule, and Dask software. We demonstrate the use of this software to run our use case and interactively visualize real-Time streams of HDF5 data.

containers↗

Clippy

Clippy (CLI + PYthon) is a Python language interface to HPC resources. Precompiled binaries that execute on HPC systems are exposed as methods to a dynamically-created Clippy Python object, where they present a familiar interface to researchers, data scientists, and others. Clippy allows these users to interact with HPC resources in an easy, straightforward environment - at the REPL, for example, or within a notebook - without the need to learn complex HPC behavior and arcane job submission commands.

Bromberger, SethA.↗

MetallData

MetallData is an HPC platform for interactive data science applications at HPC-scales. It provides an ecosystem for persistent distributed data structures, including algorithms, interactivity and storage.

Pearce, RogerA↗

In situ feature analysis for large-scale multiphase flow simulations

The study of multiphase flow is essential for designing chemical reactors such as fluidized bed reactors (FBR), as a detailed understanding of hydrodynamics is critical for optimizing reactor performance and stability. An FBR allows scientists to conduct different types of chemical reactions involving multiphase materials, especially interaction between gas and solids. During such complex chemical processes, the formation of void regions in the reactor, generally termed as bubbles, is an important phenomenon. The study of these bubbles has a deep implication in predicting the reactor’s overall efficiency. But physical experiments needed to understand bubble dynamics are costly and non-trivial due to the technical difficulties involved and harsh working conditions of the reactors. Therefore, to study such chemical processes and bubble dynamics, a state-of-the-art computational simulation MFIX-Exa is being developed. Despite the proven accuracy of MFIX-Exa in modeling bubbling phenomena, the large-scale output data prohibits the use of traditional post hoc analysis capabilities in both storage and I/O time. Herein, to address these issues and allow the application scientists to explore the bubble dynamics in an efficient and timely manner, we have developed an end-to-end analytics pipeline that enables in situ detection of bubbles, followed by a flexible post hoc visual exploration methodology of bubble dynamics. The proposed method enables interactive analysis of bubbles, along with quantification of several bubble characteristics, enabling experts to understand the bubble interactions in detail. Positive feedback from the experts has indicated the efficacy of the proposed approach for exploring bubble dynamics in very-large-scale multiphase flow simulations.

97 MATHEMATICS AND COMPUTING↗

Usage Pattern Analysis for the Summit Login Nodes

High performance computing (HPC) users interact with Summit through dedicated gateways, also known as login nodes. The performance and stability of these login nodes can have a significant impact on the user experience. In this study, the performance and stability of Summit’s five login nodes are evaluated by analyzing the log data from 2020 and 2021. The analysis focuses on the computing capability (CPU average load, users and tasks) and the storage performance, along with the associated job scheduler activity. The outcome of this study can serve as the foundation of a predictive modeling framework that enables the system admin of an HPC system to preemptively deploy countermeasures before the onset of a system failure.

Eiffert, Brett↗

Community Requirements Meta-Analysis: Characterizing Needs and Opportunities for HPDF

This High Performance Data Facility (HPDF) Project is creating a new scientific user facility to provide advanced infrastructure for data-intensive science, supporting the DOE’s Office of Science (SC) community. HPDF’s mission is to enable and accelerate scientific discovery by delivering state-of-the-art data management infrastructure, capabilities, and tools. This meta-analysis examines the needs of the breadth of the SC community, captured in publicly available community reports or mission documents. The meta-analysis identifies and provides initial characterization of fifteen core requirements for the HPDF Project team to consider during the conceptual design phase. The fifteen requirements illustrate how scientific work among SC communities requires modern, seamless user experiences across the ASCR Ecosystem to advance the use of large volumes of heterogeneous data. The scientific community requires support for the missing middle of compute between local and HPC to interactively and collaboratively use growing datasets. Data producers and end users will benefit from enhanced data catalogs and portals that improve data access through advanced search of well curated data. The fifteen requirements are examined here organized across five themes for discussion. Examples in each theme illustrate the array of scientific needs that convey the important role that the fully realized and operational High Performance Data Facility will be able to play as an integral part of the evolving ASCR Ecosystem. Our amalgamated data tables from ESnet reports demonstrate ranges to the volumes of data HPDF must be concerned with, but limitations are inherent to this meta-analysis (see Key Challenges & Limitations). Feedback and validation of these requirements along with additional details and emergent community requirements will be gathered through user research and design activities.

97 MATHEMATICS AND COMPUTING↗

Cloud Computing Methods for Near Rectilinear Halo Orbit Trajectory Design

Complicated mission design problems require innovative computational solutions. As spacecraft depart from a proposed Gateway in a Near Rectilinear Halo Orbit (NRHO), recontact analysis is required to avoid risk of collision and ensure safe operations. Escape dynamics from NRHOs are governed by multiple gravitational bodies, yielding a trajectory design space that is exhaustively large. This paper summarizes the recontact analysis for departure from the NRHO and describes how the Deep Space Trajectory Explorer (DSTE) trajectory design software incorporates high performance cloud computing to compute and visualize the orbit design space. Recent focus on exploration missions to cislunar space has kindled accelerated interest in multibody orbit solutions. Trajectory analysis in the presence of multiple gravity fields is complex, and innovative computational tools are needed to simplify complicated design spaces, to generate large quantities of data quickly, and to visualize the output for user accessibility. The Gateway mission is a prime example. The Gateway1 is proposed as a human outpost in deep space. The current baseline orbit for the Gateway is a Near Rectilinear Halo Orbit (NRHO) near the Moon.2 The NRHO exists in a regime that experiences the gravitational effects of the Earth and the Moon simultaneously, complicating orbit analysis. The mission design process benefits greatly from updated computational tools for multibody missions like the Gateway. As an example, consider the problem of assessing the risk of collision in an NRHO. As a staging location to missions to the lunar surface and beyond the Earth-Moon system, the Gateway will experience spacecraft and other objects regularly arriving and departing. Departing objects potentially include spent logistics modules, visiting crew vehicles, debris objects, wastewater particles, and cubesats. Each departure is governed by the dynamics of the Gateway orbit and the surrounding dynamical environment. Over time, any unmaintained object in such an orbit eventually departs due to the small instabilities associated with the NRHOs. A separation maneuver speeds the departure from the NRHO, but the effects of the maneuver on the spacecraft behavior depend on the location, magnitude, and direction of the burn. Escape dynamics from the NRHO with regard to these maneuver options open up an enormous potential trajectory design space where subtle changes in input can produce dramatically large changes in the results. Any departing object must avoid recontacting the Gateway as it leaves the lunar vicinity, and a recontact analysis thus involves a significant number of computations and extensive output data. To explore the dynamics of this extensive design space, the Deep Space Trajectory Explorer3 (DSTE) trajectory design software incorporates new High Performance Computing (HPC) services and novel interactive visualizations. This paper details the HPC and cloud infrastructure techniques that are implemented in the DSTE, applying the new capabilities to analysis of recontact risk with the Gateway in NRHO. NEAR RECTILINEAR HALO ORBITS The Gateway is planned to fly in a lunar NRHO as its baseline orbit. The NRHO families of orbits are subsets of the larger halo families, which originate from planar orbits near the L1 and L2 libration points; the Earth-Moon L2 halo family appears in Figure 1. Each halo orbit is perfectly periodic in the Circular Restricted 3-Body Problem (CR3BP) and becomes a quasi-periodic orbit in a higher fidelity ephemeris force model. The NRHOs are defined as those members of the halo family with bounded stability properties;2 they pass near the Moon at perilune and are nearly polar. Families exist with apolunes located both above the lunar north pole and above the lunar south pole; the Gateway is planned to reside in a southern L2 NRHO in a 9:2 resonance with the lunar synodic period. The 9:2 NRHO is characterized by a period of about 6.5 days, a perilune radius of about 3,500 km, and an apolune radius of about 71,000 km; it is strongly affected by the gravity of both the Earth and the Moon simultaneously. This NRHO offers extended communications with assets on the south pole of the Moon,4 as well as low-cost orbit maintenance and attitude control,5 favorable eclipse avoidance properties,6 and inexpensive transfers from Earth and to other destinations.5,7 The NRHO portion of the southern L2 halo family is highlighted in black in Figure 1, and the 9:2 NRHO appears in blue.

Phillips, Sean M.↗

Virtual Engineering Software Framework for Integrated Biomass Conversion Modeling

This presentation covers the design and implementation of a software tool to systematically connect computational models of unit operations to simulate an integrated process of low-temperature conversion of biomass to fuel. This virtual engineering (VE) software was designed with the overarching goal of connecting unit models written in various programming languages and requiring different computational resources within a single, flexible framework. The models and features currently considered for the VE library include mechanistic models for pretreatment, enzymatic hydrolysis, and aerobic bioreaction; high-fidelity computational fluid dynamics (CFD) simulations for enzymatic hydrolysis and aerobic bioreaction; and the capability to perform techno-economic analyses (TEA) using Aspen Plus, a commercial software package. The CFD models require access to high-performance computing (HPC) resources, so in addition to handling multiple programming languages and interfaces, the VE software must also be capable of interacting with an HPC scheduler to submit, run, and post-process jobs. Using the Python programming language, a new VE software package has been developed that contains functionality to manage the input-output communication between various unit models, schedule simulations to run on NREL's HPC and analyze those results, and interface with existing TEA software workflows. A Jupyter-notebook GUI was also created to solicit user input and provide documentation. In cases where multiple models for a particular unit-operation exist, selection between models is accomplished through a simple checkbox, with the appropriate inputs and outputs being parsed and converted seamlessly in the background. Each operation makes use of a different programming language, but the flow of information from pretreatment to enzymatic hydrolysis to bioreaction is managed with an intuitive, centralized file-communication strategy. In this talk, the programming approach and implementation details of the notebook are presented for multiple possibilities of the conversion process, including a demonstration of the ability to manage HPC resources. Additionally, an example of a sensitivity study of treatment parameters governing the overall conversion outcome is shown which highlights the ease of defining new problems using the VE Notebook workflow and leads into a discussion of ongoing work to enable outer-loop optimization studies.

biofuel↗

ChatBLAS: The First AI-Generated and Portable BLAS Library

We present ChatBLAS, the first AI-generated and portable Basic Linear Algebra Subprograms (BLAS) library on different CPU/GPU configurations. The purpose of this study is (i) to evaluate the capabilities of current large language models (LLMs) to generate a portable and HPC library for BLAS operations and (ii) to define the fundamental practices and criteria to interact with LLMs for HPC targets to elevate the trustworthiness and performance levels of the AI-generated HPC codes. The generated C/C++ codes must be highly optimized using device-specific solutions to reach high levels of performance. Additionally, these codes are very algorithm-dependent, thereby adding an extra dimension of complexity to this study. We used OpenAI’s LLM ChatGPT and focused on vector-vector BLAS level-1 operations. ChatBLAS can generate functional and correct codes, achieving high-trustworthiness levels, and can compete or even provide better performance against vendor libraries.

Valero Lara, Pedro↗

A Visual Comparison of Silent Error Propagation

High-performance computing (HPC) systems play a critical role in facilitating scientific discoveries. Their scale and complexity (e.g., the number of computational units and software stack) continue to grow as new systems are expected to process increasingly more data and reduce computing time. However, with more processing elements, the probability that these systems will experience a random bit-flip error that corrupts a program's output also increases, which is often recognized as silent data corruption. Analyzing the resiliency of HPC applications in extreme-scale computing to silent data corruption is crucial but difficult. An HPC application often contains a large number of computation units that need to be tested, and error propagation caused by error corruption is complex and difficult to interpret. Here, to accommodate this challenge, we propose an interactive visualization system that helps HPC researchers understand the resiliency of HPC applications and compare their error propagation. Our system models an application's error propagation to study a program's resiliency by constructing and visualizing its fault tolerance boundary. Coordinating with multiple interactive designs, our system enables domain experts to efficiently explore the complicated spatial and temporal correlation between error propagations. At the end, the system integrated a nonmonotonic error propagation analysis with an adjustable graph propagation visualization to help domain experts examine the details of error propagation and answer such questions as why an error is mitigated or amplified by program execution.

97 MATHEMATICS AND COMPUTING↗

Parallelizing autotuning for HPC applications: Unveiling the potential of the speculation strategy in Bayesian optimization

In the exascale computing era, tuning High-Performance Computing (HPC) applications has become a significant computational challenge. Although Bayesian optimization (BO) has emerged as a promising tool for HPC performance tuning, the BO workflow is inherently sequential (i.e., one function evaluation at a time) and cannot leverage the huge amount of parallel resources present in modern supercomputers, resulting in a considerable underutilization of their computational capabilities. This paper explores the trade-off between search quality and parallelism in BO, investigating a diverse set of methods. Building upon both previous approaches from the literature and novel methodologies introduced in this work, our study provides a deep analysis to accelerate BO performance tuning. By examining a set of synthetic functions and practical HPC applications, our exploration analyzes the interaction among various BO methods for parallelization, the quantity of parallel resources, the runtime distribution of target HPC applications, and the costs associated with different search orchestration mechanisms that have been overlooked in previous studies. Compared to sequential BO, our novel methodology achieves comparable quality while demonstrating robust scalability in search time as the amount of parallel resources increases; it also outperforms a state-of-the-art tuner, which supports parallelization, achieving up to 3.67x faster search time. We provide high-value insights for practitioners seeking to leverage the power of parallel computing for efficient HPC application tuning. Additionally, to further assist researchers in accelerating the performance tuning of their HPC applications, we provide an extension of an existing open-source tuning framework that incorporates our methods.

Bayesian optimization↗

DXT Explorer v0.1

DTX Explorer is a tool to generate interactive data visualizations of Darshan I/O traces collected from HPC applications. Its goal is to provide an easy and interactive way for researchers and developers to explore their application's I/O behavior and detect possible I/O bottlenecks that are impacting performance.

Bez, Jean Luca↗

DXT Explorer v2.0

DTX Explorer is a tool to generate interactive data visualizations of Darshan I/O traces collected from HPC applications. Its goal is to provide an easy and interactive way for researchers and developers to explore their application's I/O behavior and detect possible I/O bottlenecks that are impacting performance.

Bez, JeanLuca↗

Integrating Artificial Intelligence into Science Gateways

Science gateways are altering the manner in which people interact with high performance computing (HPC) by providing a web browser based interface to advanced computing platforms. In particular, science gateways lower the barrier to using HPC by simplifying the process of submitting workloads to such systems and by offloading the efforts required to use HPC to the maintainers of the system. While science gateways decrease the time-to-science that comes with using such advanced systems, progress can still be made in improving the user's experience. In this paper we explore two strategies for integrating artificial intelligence tools commonly found in non-HPC service workflows: voice activated assistants and chatbots. Since August 2021, the HPC group at Idaho National Laboratory answers an average of 581 support tickets per month of which a large percentage could be addressed via these two strategies. This work defines the key capabilities that an HPC voice activated assistant and chatbot would need to address for a userbase consisting of largely non-expert users as well as a design for integration into the Open OnDemand science gateway.

97 MATHEMATICS AND COMPUTING↗

Understanding Interactive and Reproducible Computing With Jupyter Tools at Facilities

Increasingly Jupyter tools are being adopted and incorporated into High Performance Computing (HPC) and scientific user facilities. Adopting Jupyter tools enables more interactive and reproducible computational work at facilities across data life cycles. As the volume, variety, and scope of data grow, scientists need to be able to analyze and share results in user friendly ways. Human-centered research highlights design challenges around computational notebooks, and our qualitative user study shifts focus to better characterize how Jupyter tools are being used in HPC and science user facilities today. We conducted twenty-nine interviews, and obtained 103 survey responses from NERSC Jupyter users, to better understand the increasing role of interactive computing tools in DOE sponsored scientific work. We examine a range of issues that emerge using and supporting Jupyter in HPC ecosystems, including: how Jupyter is being used by scientists in HPC and user facility ecosystems; how facilities are purposefully supporting Jupyter in their ecosystems; feedback NERSC users have about the facility’s deployment, and, discuss features NERSC indicated would be helpful. We offer a variety of takeaways for staff supporting Jupyter at facilities, Project Jupyter and related open source communities, and funding agencies supporting interactive computing work.

97 MATHEMATICS AND COMPUTING↗

VISILIENCE: An Interactive Visualization Framework for Resilience Analysis using Control-Flow Graph

Soft errors have become one of the major concerns for the error resilience of HPC applications, as those errors can cause HPC applications to generate serious outcomes such as Silent Data Corruptions (SDCs). A large body of approaches has been proposed to analyze the resilience of HPC applications. However, existing studies rarely address the challenges of the analysis result perception. Specifically, resilience analysis techniques often produce a massive volume of unstructured data, making it difficult for programmers to conduct the resilience analysis due to non-intuitive raw data. Furthermore, different analysis models produce diverse results with multiple levels of details, which may create hurdles to compare and explore the resilience of HPC program execution. To this end, we present VISILIENCE, an interactive VISual resILIENCE analysis framework to allow programmers to facilitate the resilience analysis of HPC applications. In particular, VISILIENCE leverages an effective visualization approach Control Flow Graph (CFG) to present a function execution. In addition, three widely-used models for resilience analysis (i.e., Y-Branch, IPAS, and TRIDENT) are seamlessly embedded into the framework for resilience analysis and result comparison. Multiple case studies have been conducted to demonstrate the effectiveness of our proposed framework VISILIENCE.

Jiang, Hailong↗