Search NASA⌕ Search

SEARCH · Search NASA

Results for “Performance Trace”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Chimbuko: A Workflow-Level Scalable Performance Trace Analysis Tool

ABSTRACT Due to the sheer volume of data it is typically impractical to analyze the detailed performance of an HPC application running at-scale. While conventional small-scale benchmarking and scaling studies are often sufficient for simple applications, many modern workflow-based applications couple multiple elements with competing resource demands and complex inter-communication patterns for which performance cannot easily be studied in isolation and at small scale. This work discusses Chimbuko, a performance analysis framework that provides real-time, in situ anomaly detection. By focusing specifically on performance anomalies and their origin (aka provenance), data volumes are dramatically reduced without losing necessary details. To the best of our knowledge, Chimbuko is the first online, distributed, and scalable workflow-level performance trace analysis framework. We demonstrate the tool's usefulness on Oak Ridge National Laboratory's Summit system.

97 MATHEMATICS AND COMPUTING↗

Chimbuko: A Workflow-Level Scalable Performance Trace Analysis Tool

Due to the sheer volume of data it is typically impractical to analyze the detailed performance of an HPC application running at-scale. While conventional small-scale benchmarking and scaling studies are often sufficient for simple applications, many modern workflow-based applications couple multiple elements with competing resource demands and complex inter-communication patterns for which performance cannot easily be studied in isolation and at small scale. This work discusses Chimbuko, a performance analysis framework that provides real-time, in situ anomaly detection. By focusing specifically on performance anomalies and their origin (aka provenance), data volumes are dramatically reduced without losing necessary details. To the best of our knowledge, Chimbuko is the first online, distributed, and scalable workflow-level performance trace analysis framework. We demonstrate the tool's usefulness on Oak Ridge National Laboratory's Summit system.

97 MATHEMATICS AND COMPUTING↗

Traveler: Navigating Task Parallel Traces for Performance Analysis

Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.

97 MATHEMATICS AND COMPUTING↗

Signal Processing Based Method for Real-Time Anomaly Detection in High-Performance Computing

Performance anomalies can manifest as irregular execution times or abnormal execution events for many reasons, including network congestion and resource contention. Detecting such anomalies in real-time by analyzing the details of performance traces at scale is impractical due to the sheer volume of data High-Performance Computing (HPC) applications produce. In this paper, we propose formulating HPC performance anomaly detection as a signal-processing problem where anomalies can be treated as noise. We evaluate our proposed method in comparison with two other commonly used anomaly detection techniques of varying complexity based on their detection accuracy and scalability. Since real-time in-situ anomaly detection at a large scale requires lightweight methods that can handle a large volume of streaming data, we find that our proposed method provides the best trade-off. We then implement the proposed method in Chimbuko, the first online, distributed, and scalable workflow-level performance trace analysis framework. We compare our proposed signal-based anomaly detection algorithm with two other methods using a function of their accuracy, F1 score, and detection overhead. Our experiments demonstrate that our proposed approach achieves a 99% improvement for the benchmark datasets and a 93% improvement with Chimbuko traces.

99 GENERAL AND MISCELLANEOUS↗

Characterization of inexpensive metal oxide sensor performance for trace methane detection

Abstract. Methane, a major contributor to climate change, is emitted by a variety of natural and anthropogenic sources. Commercially available lab-grade instruments for sensing trace methane are expensive, and previous efforts to develop inexpensive, field-deployable trace methane sensors have had mixed results. Industrial and commercial metal oxide (MOx) methane sensors, which are intended for leak detection and safety monitoring, can potentially be repurposed and adapted for low-concentration sensing. As an initial step towards developing a low-cost sensing system, we characterize the performance of five off-the-shelf MOx sensors for 2–10 ppm methane detection in a laboratory setting (Figaro Engineering TGS2600, TGS2602, TGS2611-C00, TGS2611-E00, and Henan Hanwei Electronics MQ4). We identify TGS2611-C00, TGS2611-E00, and MQ4 as promising for trace methane sensing but show that variations in ambient humidity and temperature pose a challenge for the sensors in this application.

54 ENVIRONMENTAL SCIENCES↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Computational Performance Bounds Prediction in Quantum Computing With Unstable Noise

Quantum computing has significantly advanced in recent years, boasting devices with hundreds of quantum bits (qubits), hinting at its potential quantum advantage over classical computing. Yet, noise in quantum devices poses significant barriers to realizing this supremacy. Understanding noise’s impact is crucial for reproducibility and application reuse; moreover, the next-generation quantum-centric supercomputing essentially requires efficient and accurate noise characterization to support system management (e.g., job scheduling), where ensuring correct functional performance (i.e., fidelity) of jobs on available quantum devices can even be higher-priority than traditional objectives. However, noise fluctuates over time, even on the same quantum device, which makes predicting the computational bounds for on-the-fly noise is vital. Noisy quantum simulation can offer insights but faces efficiency and scalability issues. Here, in this work, we propose a data-driven workflow, namely QuBound, to predict computational performance bounds. It decomposes historical performance traces to isolate noise sources and devises a novel encoder to embed circuit and noise information processed by a Long Short-Term Memory (LSTM) network. For evaluation, we compare QuBound with a state-of-the-art learning-based predictor, which only generates a single performance value instead of a bound. Experimental results show that the result of the existing approach falls outside of performance bounds, while all predictions from our QuBound with the assistance of performance decomposition better fit the bounds. Moreover, QuBound can efficiently produce practical bounds for various circuits with over 106 speedup over simulation; in addition, the range from QuBound is over 10× narrower than the state-of-the-art analytical approach.

Li, Jinyang [George Mason Univ., Fairfax, VA (Unit↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Panorama 360 (Final Report)

This final technical report from the lead institution, USC grant #DE-SC0012636, serves as the final technical report for collaborative institution UNC-CH grant #DE-SC0012390. The goal was to develop a repository and associated capabilities for data collection, ingestion, and analysis for a broad class of DOE applications that span experimental and simulation science workflows. In particular, this work focuses on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: (1) A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); (2) A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; (3) A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and (4) Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Composition, Emissions, and Air Quality Impacts of Hazardous Air Pollutants in Unburned Natural Gas from Residential Stoves in California

The presence of hazardous air pollutants (HAPs) entrained in end-use natural gas (NG) is an understudied source of human health risks. We performed trace gas analyses on 185 unburned NG samples collected from 159 unique residential NG stoves across seven geographic regions in California. Our analyses commonly detected 12 HAPs with significant variability across region and gas utility. Mean regional benzene, toluene, ethylbenzene, and total xylenes (BTEX) concentrations in end-use NG ranged from 1.6–25 ppmv–benzene alone was detected in 99% of samples, and mean concentrations ranged from 0.7–12 ppmv (max: 66 ppmv). By applying previously reported NG and methane emission rates throughout California’s transmission, storage, and distribution systems, we estimated statewide benzene emissions of 4,200 (95% CI: 1,800–9,700) kg yr –1 that are currently not included in any statewide inventories–equal to the annual benzene emissions from nearly 60,000 light-duty gasoline vehicles. Additionally, we found that NG leakage from stoves and ovens while not in use can result in indoor benzene concentrations that can exceed the California Office of Environmental Health Hazard Assessment 8-h Reference Exposure Level of 0.94 ppbv–benzene concentrations comparable to environmental tobacco smoke. This study supports the need to further improve our understanding of leaked downstream NG as a source of health risk.

42 ENGINEERING↗

A new metrics framework for quantifying and intercomparing atmospheric rivers in observations, reanalyses, and climate models

We present a new atmospheric river (AR) analysis and benchmarking tool, namely Atmospheric River Metrics Package (ARMP). It includes a suite of new AR metrics that are designed for quick analysis of AR characteristics via statistics in gridded climate datasets such as model output and reanalysis. This package can be used for climate model evaluation in comparison with reanalysis and observational products. Integrated metrics such as mean bias and spatial pattern correlation are efficient for diagnosing systematic AR biases in climate models. For example, the package identifies the fact that, in CMIP5 and CMIP6 (Coupled Model Intercomparison Project Phases 5 and 6) models, AR tracks in the South Atlantic are positioned farther poleward compared to ERA5 reanalysis, while in the South Pacific, tracks are generally biased towards the Equator. For the landfalling AR peak season, we find that most climate models simulate a completely opposite seasonal cycle over western Africa. This tool can also be used for identifying and characterizing structural differences among different AR detectors (ARDTs). For example, ARs detected with the Mundhenk algorithm exhibit systematically larger size, width, and length compared to the TempestExtremes (TE) method. The AR metrics developed from this work can be routinely applied for model benchmarking and during the development cycle to trace performance evolution across model versions or generations and set objective targets for the improvement of models. They can also be used by operational centers to perform near-real-time climate and extreme event impact assessments as part of their forecast cycle.

58 GEOSCIENCES↗

High Performance Computing Application I/O Traces

The dataset comprises trace files from high performance computing (HPC) simulations. The trace files contain records of every I/O operation executed by a simulation application run, including I/O operations from HDF5, MPI-IO, and POSIX and all of the parameters supplied to those operations, e.g. file name, offset, and flags. The traces are generated by executing a simulation application that is linked with the Recorder tracing tool (https://github.com/uiuc-hpc/Recorder). The Recorder trace tool intercepts the I/O calls made by the application, records the I/O trace record, and then calls the intended I/O call so that the operation executes.

Wang, C.↗

Quantum information approach to high energy interactions

High energy hadron interactions are commonly described by using a probabilistic parton model that ignores quantum entanglement present in the light-cone wave functions. Here, we argue that since a high energy interaction samples an instant snapshot of the hadron wave function, the phases of different Fock state wave functions cannot be measured—therefore the light-cone density matrix has to be traced over these unobservable phases. Performing this trace with the corresponding U(1) Haar integration measure leads to ‘Haar scrambling’ of the density matrix, and to the emergence of entanglement entropy. This entanglement entropy is determined by the Fock state probability distribution, and is thus directly related to the parton structure functions. As proposed earlier, at large rapidity η the hadron state becomes maximally entangled, and the entanglement entropy is S E ~η according to QCD evolution equations. When the phases of Fock state components are controlled, for example in spin asymmetry measurements, the Haar average cannot be performed, and the probabilistic parton description breaks down.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Turbulent impurity transport simulations in Wendelstein 7-X plasmas

A study of turbulent impurity transport by means of quasilinear and nonlinear gyrokinetic simulations is presented for Wendelstein 7-X (W7-X). The calculations have been carried out with the recently developed gyrokinetic code stella. Different impurity species are considered in the presence of various types of background instabilities: ion temperature gradient (ITG), trapped electron mode (TEM) and electron temperature gradient (ETG) modes for the quasilinear part of the work; ITG and TEM for the nonlinear results. While the quasilinear approach allows one to draw qualitative conclusions about the sign or relative importance of the various contributions to the flux, the nonlinear simulations quantitatively determine the size of the turbulent flux and check the extent to which the quasilinear conclusions hold. Although the bulk of the nonlinear simulations are performed at trace impurity concentration, nonlinear simulations are also carried out at realistic effective charge values, in order to know to what degree the conclusions based on the simulations performed for trace impurities can be extrapolated to realistic impurity concentrations. The presented results conclude that the turbulent radial impurity transport in W7-X is mainly dominated by ordinary diffusion, which is close to that measured during the recent W7-X experimental campaigns. Finally, it is also confirmed that thermodiffusion adds a weak inward flux contribution and that, in the absence of impurity temperature and density gradients, ITG- and TEM-driven turbulence push the impurities inwards and outwards, respectively.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Optimization of ray-tracing simulations to confirm performance of the GP-SANS instrument at the High-Flux Isotope Reactor

The CG-2 beamline at the High Flux Isotope Reactor (HFIR) exhibits a notable discrepancy between observed count rates and the count rates we would expect based on a Monte-Carlo neutron ray-trace simulation. These simulations consistently predict count rates approximately five times greater than those observed in four separate experimental runs involving different instrument configurations. This discrepancy suggests that certain factors are causing losses in measurements that are not adequately accounted for in the simulation, in particular guide reflectivity or misalignment. To investigate these discrepancies, a high-dimensional simulation parameter approach is applied in order to understand the losses. Region of Interest (ROI) groups along the instrument are assigned to different surfaces of the guide components within the simulation. This allows the parameters of those guide components to be varied as a group to minimize the complexity of the search space. The result is an optimization of simulation parameters using an iterative scheme that aims to minimize the difference between experimentally measured count rates and simulated count rates across all tested collimator combinations. This proposed methodology holds the potential to reveal previously unrecognized sources of intensity loss in the CG-2 beamline at HFIR and improve the accuracy of simulations, leading to enhanced understanding and performance of the beamline for various scientific applications.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Data for NB6 HBRR Science Design ORNL/TM-2025/3807

Data for the report (ORNL/TM-2025/3807) that describes the calculations and the Monte Carlo Ray Tracing simulations performed using the McStas package to determine the coatings and geometry for the NB-6 guide. It provides the information to inform the mechanical design, validation tests and verification that it meets the science requirements.

47 OTHER INSTRUMENTATION↗

Ray-Tracing Simulations Characterising the Performance of the Proposed 2024 HFIR HB4 Main Shutter

The Main Shutter at HB4 will serve two purposes after the HFIR Beryllium Reflector Replacement planned to take place in 2024. First as the primary certified safety control controlling the passage of neutrons from the cold source in the HFIR pressure vessel into the cold guide hall, and second as the first set of reflecting surfaces used to guide neutrons from the source and into the individual guide starts for each instrument in the cold guide hall.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗