Search NASA⌕ Search

SEARCH · Search NASA

Results for “Darshan log”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

April 2020 Darshan counters from the Summit supercomputer

This dataset is the Darshan counters collected from the Summit supercomputer in a month of April 2020. 1. Description of methods used for collection/generation of data: Job submitted on Summit HPC system when completed successfully and has made I/O calls (captured by Darshan tool) writes a Darshan log file on alpine filesystem. One job can have multiple `jsrun` commands and Darshan will generate separate logs each log corresponding to an `jsrun` command, so a job can have one or more Darshan logs associated with it. 2. Methods for processing the data: To process the data, we first use `darshan-util` tool to parse the Darshan logs. Then we restructure the logs and merge data from multiple Darshan logs if they belong to the same Summit job.

97 MATHEMATICS AND COMPUTING↗

Summit Darshan Archival Dataset

Summit Darshan Archival Dataset contains 2021 Summit Darshan log data for 25 applications and is grouped into science domains. The dataset is processed, and all the propriety fields are anonymized. The resultant data is converted into a tabular structure and saved in parquet file format. In this notebook, we demonstrate how to access the data. Data Organization: The data is organized into two directories: Darshan total (`darshan_total`): List all the high levels generated by the `darshan-parser --total` command on `.darshan` files. There is one parquet file for each application. Note: `uid` and `exe` field are masked Darshan detail (`darshan_detail`): This data contains detailed job level log information extracted by command `darshan-parser` on the raw `.darshan` files. The data is sorted by directory hierarchy in the order of `year/month/day (2021/12/07)`. For instance, to get the data for a `job_id` 3819766 of application `App11`, which was executed on `2021-12-07`can be accessed as follows. Note:`uid` and `filename` fields are masked

97 MATHEMATICS AND COMPUTING↗

Characterizing Machine Learning I/O Workloads on Leadership Scale HPC Systems

High performance computing (HPC) is no longer solely limited to traditional workloads such as simulation and modeling. With the increase in the popularity of machine learning (ML) and deep learning (DL) technologies, we are observing that an increasing number of HPC users are incorporating ML methods into their workflow and scientific discovery processes, across a wide spectrum of science domains such as biology, earth science, and physics. This gives rise to a diverse set of I/O patterns than the traditional checkpoint/restart-based HPC I/O behavior. The details of the I/O characteristics of such ML I/O workloads have not been studied extensively for large-scale leadership HPC systems. This paper aims to fill that gap by providing an in-depth analysis to gain an understanding of the I/O behavior of ML I/O workloads using darshan - an I/O characterization tool designed for lightweight tracing and profiling. We study the darshan logs of more than 23, 000 HPC ML I/O jobs over a time period of one year running on Summit - the second-fastest supercomputer in the world. This paper provides a systematic I/O characterization of ML I/O jobs running on a leadership scale supercomputer to understand how the I/O behavior differs across science domains and the scale of workloads, and analyze the usage of parallel file system and burst buffer by ML I/O workloads.

Paul, Arnab↗

Darshan for HEP applications

Modern HEP workflows must manage increasingly large and complex data collections. HPC facilities may be employed to help meet these workflows’ growing data processing needs. However, a better understanding of the I/O patterns and underlying bottlenecks of these workflows is necessary to meet the performance expectations of HPC systems.Darshan is a lightweight I/O characterization tool that captures concise views of HPC application I/O behavior. It intercepts application I/O calls at runtime, records file access statistics for each process, and generates log files detailing application I/O access patterns.Typical HEP workflows include event generation, detector simulation, event reconstruction, and subsequent analysis stages. A study of the I/O behavior of the ATLAS simulation and filtering stage, and the CMS simulation workflow using Darshan is presented, including insights into the I/O operations and data access size.

Wang, Rui↗

I/O Bottleneck Detection and Tuning: Connecting the Dots using Interactive Log Analysis

Using parallel file systems efficiently is a tricky problem due to inter-dependencies among multiple layers of I/O software, including high-level I/O libraries (HDF5, netCDF, etc.), MPI-IO, POSIX, and file systems (GPFS, Lustre, etc.). Profiling tools such as Darshan collect traces to help understand the I/O performance behavior. However, there are significant gaps in analyzing the collected traces and then applying tuning options offered by various layers of I/O software. Seeking to connect the dots between I/O bottleneck detection and tuning, we propose DXT Explorer, an interactive log analysis tool. In this paper, we present a case study using our interactive log analysis tool to identify and apply various I/O optimizations. We report an evaluation of performance improvement achieved for four I/O kernels extracted from science applications.

Bez, Jean Luca↗

BASS. XXV. DR2 Broad-line-based Black Hole Mass Estimates and Biases from Obscuration

We present measurements of broad emission lines and virial estimates of supermassive black hole masses (M BH ) for a large sample of ultrahard X-ray-selected active galactic nuclei (AGNs) as part of the second data release of the BAT AGN Spectroscopic Survey (BASS/DR2). Our catalog includes M BH estimates for a total of 689 AGNs, determined from the Hα, Hβ, Mg II λ2798, and/or C IV λ1549 broad emission lines. The core sample includes a total of 512 AGNs drawn from the 70 month Swift/BAT all-sky catalog. We also provide measurements for 177 additional AGNs that are drawn from deeper Swift/BAT survey data. We study the links between M BH estimates and line-of-sight obscuration measured from X-ray spectral analysis. We find that broad Hα emission lines in obscured AGNs ($\mathrm{log}({N}_{{\rm{H}}}/{\mathrm{cm}}^{-2})\gt 22.0$) are on average a factor of ${8.0}_{-2.4}^{+4.1}$ weaker relative to ultrahard X-ray emission and about ${35}_{-12}^{\,+7}$% narrower than those in unobscured sources (i.e., $\mathrm{log}({N}_{{\rm{H}}}/{\mathrm{cm}}^{-2})\lt 21.5$). This indicates that the innermost part of the broad-line region is preferentially absorbed. Consequently, current single-epoch M BH prescriptions result in severely underestimated (>1 dex) masses for Type 1.9 sources (AGNs with broad Hα but no broad Hβ) and/or sources with $\mathrm{log}({N}_{{\rm{H}}}/{\mathrm{cm}}^{-2})\gtrsim 22.0$. We provide simple multiplicative corrections for the observed luminosity and width of the broad Hα component (L[bHα] and FWHM[bHα]) in such sources to account for this effect and to (partially) remedy M BH estimates for Type 1.9 objects. As a key ingredient of BASS/DR2, our work provides the community with the data needed to further study powerful AGNs in the low-redshift universe.

79 ASTRONOMY AND ASTROPHYSICS↗

BASS XXXII: Studying the Nuclear Millimeter-wave Continuum Emission of AGNs with ALMA at Scales ≲100–200 pc

To understand the origin of nuclear (≲100 pc) millimeter-wave (mm-wave) continuum emission in active galactic nuclei (AGNs), we systematically analyzed subarcsecond resolution Band-6 (211–275 GHz) Atacama Large Millimeter/submillimeter Array data of 98 nearby AGNs (z < 0.05) from the 70 month Swift/BAT catalog. The sample, almost unbiased for obscured systems, provides the largest number of AGNs to date with high mm-wave spatial resolution sampling (~1–200 pc), and spans broad ranges of 14–150 keV luminosity {$40\lt \mathrm{log}[{L}_{14-150}/(\mathrm{erg}\,{{\rm{s}}}^{-1})]\lt 45$}, black hole mass $[5\lt \mathrm{log}({M}_{\mathrm{BH}}/{M}_{\odot })\lt 10$], and Eddington ratio ($-4\lt \mathrm{log}{\lambda }_{\mathrm{Edd}}\lt 2$). We find a significant correlation between 1.3 mm (230 GHz) and 14–150 keV luminosities. Its scatter is ≈0.36 dex, and the mm-wave emission may serve as a good proxy of the AGN luminosity, free of dust extinction up to N H ~ 10 26 cm –2 . While the mm-wave emission could be self-absorbed synchrotron radiation around the X-ray corona according to past works, we also discuss different possible origins of the mm-wave emission: AGN-related dust emission, outflow-driven shocks, and a small-scale (<200 pc) jet. The dust emission is unlikely to be dominant, as the mm-wave slope is generally flatter than expected. Also, due to no increase in the mm-wave luminosity with the Eddington ratio, a radiation-driven outflow model is possibly not the common mechanism. Furthermore, we find independence of the mm-wave luminosity on indicators of the inclination angle from the polar axis of the nuclear structure, which is inconsistent with a jet model whose luminosity depends only on the angle.

79 ASTRONOMY AND ASTROPHYSICS↗