Search NASA⌕ Search

SEARCH · Search NASA

Results for “Performance Trace”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Chimbuko: A Workflow-Level Scalable Performance Trace Analysis Tool

ABSTRACT Due to the sheer volume of data it is typically impractical to analyze the detailed performance of an HPC application running at-scale. While conventional small-scale benchmarking and scaling studies are often sufficient for simple applications, many modern workflow-based applications couple multiple elements with competing resource demands and complex inter-communication patterns for which performance cannot easily be studied in isolation and at small scale. This work discusses Chimbuko, a performance analysis framework that provides real-time, in situ anomaly detection. By focusing specifically on performance anomalies and their origin (aka provenance), data volumes are dramatically reduced without losing necessary details. To the best of our knowledge, Chimbuko is the first online, distributed, and scalable workflow-level performance trace analysis framework. We demonstrate the tool's usefulness on Oak Ridge National Laboratory's Summit system.

97 MATHEMATICS AND COMPUTING↗

Chimbuko: A Workflow-Level Scalable Performance Trace Analysis Tool

Due to the sheer volume of data it is typically impractical to analyze the detailed performance of an HPC application running at-scale. While conventional small-scale benchmarking and scaling studies are often sufficient for simple applications, many modern workflow-based applications couple multiple elements with competing resource demands and complex inter-communication patterns for which performance cannot easily be studied in isolation and at small scale. This work discusses Chimbuko, a performance analysis framework that provides real-time, in situ anomaly detection. By focusing specifically on performance anomalies and their origin (aka provenance), data volumes are dramatically reduced without losing necessary details. To the best of our knowledge, Chimbuko is the first online, distributed, and scalable workflow-level performance trace analysis framework. We demonstrate the tool's usefulness on Oak Ridge National Laboratory's Summit system.

97 MATHEMATICS AND COMPUTING↗

Traveler: Navigating Task Parallel Traces for Performance Analysis

Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.

97 MATHEMATICS AND COMPUTING↗

Signal Processing Based Method for Real-Time Anomaly Detection in High-Performance Computing

Performance anomalies can manifest as irregular execution times or abnormal execution events for many reasons, including network congestion and resource contention. Detecting such anomalies in real-time by analyzing the details of performance traces at scale is impractical due to the sheer volume of data High-Performance Computing (HPC) applications produce. In this paper, we propose formulating HPC performance anomaly detection as a signal-processing problem where anomalies can be treated as noise. We evaluate our proposed method in comparison with two other commonly used anomaly detection techniques of varying complexity based on their detection accuracy and scalability. Since real-time in-situ anomaly detection at a large scale requires lightweight methods that can handle a large volume of streaming data, we find that our proposed method provides the best trade-off. We then implement the proposed method in Chimbuko, the first online, distributed, and scalable workflow-level performance trace analysis framework. We compare our proposed signal-based anomaly detection algorithm with two other methods using a function of their accuracy, F1 score, and detection overhead. Our experiments demonstrate that our proposed approach achieves a 99% improvement for the benchmark datasets and a 93% improvement with Chimbuko traces.

99 GENERAL AND MISCELLANEOUS↗

Characterization of inexpensive metal oxide sensor performance for trace methane detection

Abstract. Methane, a major contributor to climate change, is emitted by a variety of natural and anthropogenic sources. Commercially available lab-grade instruments for sensing trace methane are expensive, and previous efforts to develop inexpensive, field-deployable trace methane sensors have had mixed results. Industrial and commercial metal oxide (MOx) methane sensors, which are intended for leak detection and safety monitoring, can potentially be repurposed and adapted for low-concentration sensing. As an initial step towards developing a low-cost sensing system, we characterize the performance of five off-the-shelf MOx sensors for 2–10 ppm methane detection in a laboratory setting (Figaro Engineering TGS2600, TGS2602, TGS2611-C00, TGS2611-E00, and Henan Hanwei Electronics MQ4). We identify TGS2611-C00, TGS2611-E00, and MQ4 as promising for trace methane sensing but show that variations in ambient humidity and temperature pose a challenge for the sensors in this application.

54 ENVIRONMENTAL SCIENCES↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Analysis of hypercube cache performance using address traces generated by TRAPEDS

The authors utilize a recently developed software method of capturing and analyzing address traces, known as TRAPEDS (TRAce Producing Execution Driven Simulation), to provide address traces for cache performance evaluation on a hypercube multicomputer. Utilizing TRAPEDS user code traces obtained from the implementations of several parallel algorithms on the Intel iPSC/2 hypercube, the authors simulate the cache performance effects of changing cache size, line size, and set associativity. Particular attention is devoted to the effect on cache performance of changing the dimension of the hypercube for a particular program and to the variation in cache statistics among the nodes of the hypercube.

Stunkel, Craig B.↗

Computational Performance Bounds Prediction in Quantum Computing With Unstable Noise

Quantum computing has significantly advanced in recent years, boasting devices with hundreds of quantum bits (qubits), hinting at its potential quantum advantage over classical computing. Yet, noise in quantum devices poses significant barriers to realizing this supremacy. Understanding noise’s impact is crucial for reproducibility and application reuse; moreover, the next-generation quantum-centric supercomputing essentially requires efficient and accurate noise characterization to support system management (e.g., job scheduling), where ensuring correct functional performance (i.e., fidelity) of jobs on available quantum devices can even be higher-priority than traditional objectives. However, noise fluctuates over time, even on the same quantum device, which makes predicting the computational bounds for on-the-fly noise is vital. Noisy quantum simulation can offer insights but faces efficiency and scalability issues. Here, in this work, we propose a data-driven workflow, namely QuBound, to predict computational performance bounds. It decomposes historical performance traces to isolate noise sources and devises a novel encoder to embed circuit and noise information processed by a Long Short-Term Memory (LSTM) network. For evaluation, we compare QuBound with a state-of-the-art learning-based predictor, which only generates a single performance value instead of a bound. Experimental results show that the result of the existing approach falls outside of performance bounds, while all predictions from our QuBound with the assistance of performance decomposition better fit the bounds. Moreover, QuBound can efficiently produce practical bounds for various circuits with over 106 speedup over simulation; in addition, the range from QuBound is over 10× narrower than the state-of-the-art analytical approach.

Li, Jinyang [George Mason Univ., Fairfax, VA (Unit↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Panorama 360 (Final Report)

This final technical report from the lead institution, USC grant #DE-SC0012636, serves as the final technical report for collaborative institution UNC-CH grant #DE-SC0012390. The goal was to develop a repository and associated capabilities for data collection, ingestion, and analysis for a broad class of DOE applications that span experimental and simulation science workflows. In particular, this work focuses on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: (1) A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); (2) A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; (3) A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and (4) Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Composition, Emissions, and Air Quality Impacts of Hazardous Air Pollutants in Unburned Natural Gas from Residential Stoves in California

The presence of hazardous air pollutants (HAPs) entrained in end-use natural gas (NG) is an understudied source of human health risks. We performed trace gas analyses on 185 unburned NG samples collected from 159 unique residential NG stoves across seven geographic regions in California. Our analyses commonly detected 12 HAPs with significant variability across region and gas utility. Mean regional benzene, toluene, ethylbenzene, and total xylenes (BTEX) concentrations in end-use NG ranged from 1.6–25 ppmv–benzene alone was detected in 99% of samples, and mean concentrations ranged from 0.7–12 ppmv (max: 66 ppmv). By applying previously reported NG and methane emission rates throughout California’s transmission, storage, and distribution systems, we estimated statewide benzene emissions of 4,200 (95% CI: 1,800–9,700) kg yr –1 that are currently not included in any statewide inventories–equal to the annual benzene emissions from nearly 60,000 light-duty gasoline vehicles. Additionally, we found that NG leakage from stoves and ovens while not in use can result in indoor benzene concentrations that can exceed the California Office of Environmental Health Hazard Assessment 8-h Reference Exposure Level of 0.94 ppbv–benzene concentrations comparable to environmental tobacco smoke. This study supports the need to further improve our understanding of leaked downstream NG as a source of health risk.

42 ENGINEERING↗

A new metrics framework for quantifying and intercomparing atmospheric rivers in observations, reanalyses, and climate models

We present a new atmospheric river (AR) analysis and benchmarking tool, namely Atmospheric River Metrics Package (ARMP). It includes a suite of new AR metrics that are designed for quick analysis of AR characteristics via statistics in gridded climate datasets such as model output and reanalysis. This package can be used for climate model evaluation in comparison with reanalysis and observational products. Integrated metrics such as mean bias and spatial pattern correlation are efficient for diagnosing systematic AR biases in climate models. For example, the package identifies the fact that, in CMIP5 and CMIP6 (Coupled Model Intercomparison Project Phases 5 and 6) models, AR tracks in the South Atlantic are positioned farther poleward compared to ERA5 reanalysis, while in the South Pacific, tracks are generally biased towards the Equator. For the landfalling AR peak season, we find that most climate models simulate a completely opposite seasonal cycle over western Africa. This tool can also be used for identifying and characterizing structural differences among different AR detectors (ARDTs). For example, ARs detected with the Mundhenk algorithm exhibit systematically larger size, width, and length compared to the TempestExtremes (TE) method. The AR metrics developed from this work can be routinely applied for model benchmarking and during the development cycle to trace performance evolution across model versions or generations and set objective targets for the improvement of models. They can also be used by operational centers to perform near-real-time climate and extreme event impact assessments as part of their forecast cycle.

58 GEOSCIENCES↗

Scalable Performance Environments for Parallel Systems

As parallel systems expand in size and complexity, the absence of performance tools for these parallel systems exacerbates the already difficult problems of application program and system software performance tuning. Moreover, given the pace of technological change, we can no longer afford to develop ad hoc, one-of-a-kind performance instrumentation software; we need scalable, portable performance analysis tools. We describe an environment prototype based on the lessons learned from two previous generations of performance data analysis software. Our environment prototype contains a set of performance data transformation modules that can be interconnected in user-specified ways. It is the responsibility of the environment infrastructure to hide details of module interconnection and data sharing. The environment is written in C++ with the graphical displays based on X windows and the Motif toolkit. It allows users to interconnect and configure modules graphically to form an acyclic, directed data analysis graph. Performance trace data are represented in a self-documenting stream format that includes internal definitions of data types, sizes, and names. The environment prototype supports the use of head-mounted displays and sonic data presentation in addition to the traditional use of visual techniques.

Reed, Daniel A.↗

Enhanced Performance of Streamline-Traced External-Compression Supersonic Inlets

A computational design study was conducted to enhance the aerodynamic performance of streamline-traced, external-compression inlets for Mach 1.6. Compared to traditional external-compression, two-dimensional and axisymmetric inlets, streamline-traced inlets promise reduced cowl wave drag and sonic boom, but at the expense of reduced total pressure recovery and increased total pressure distortion. The current study explored a new parent flowfield for the streamline tracing and several variations of inlet design factors, including the axial displacement and angle of the subsonic cowl lip, the vertical placement of the engine axis, and the use of porous bleed in the subsonic diffuser. The performance was enhanced over that of an earlier streamline-traced inlet such as to increase the total pressure recovery and reduce total pressure distortion.

performance↗

Enhanced Performance of Streamline-Traced External-Compression Supersonic Inlets

A computational design study was conducted to enhance the aerodynamic performance of streamline-traced, external-compression inlets for Mach 1.6. The current study explored a new parent flowfield for the streamline tracing and several variations of inlet design factors, including the axial displacement and angle of the subsonic cowl lip, the vertical placement of the engine axis, and the use of porous bleed in the subsonic diffuser. The performance was enhanced over that of an earlier streamline-traced inlet such as to increase the total pressure recovery and reduce total pressure distortion

performance↗

Performance Testing of a Trace Contaminant Control Subassembly for the International Space Station

As part of the International Space Station (ISS) Trace Contaminant Control Subassembly (TCCS) development, a performance test has been conducted to provide reference data for flight verification analyses. This test, which used the U.S. Habitation Module (U.S. Hab) TCCS as the test article, was designed to add to the existing database on TCCS performance. Included in this database are results obtained during ISS development testing; testing of functionally similar TCCS prototype units; and bench scale testing of activated charcoal, oxidation catalyst, and granular lithium hydroxide (LiOH). The present database has served as the basis for the development and validation of a computerized TCCS process simulation model. This model serves as the primary means for verifying the ISS TCCS performance. In order to mitigate risk associated with this verification approach, the U.S. Hab TCCS performance test provides an additional set of data which serve to anchor both the process model and previously-obtained development test data to flight hardware performance. The following discussion provides relevant background followed by a summary of the test hardware, objectives, requirements, and facilities. Facility and test article performance during the test is summarized, test results are presented, and the TCCS's performance relative to past test experience is discussed. Performance predictions made with the TCCS process model are compared with the U.S. Hab TCCS test results to demonstrate its validation.

Perry, J. L.↗