Search NASA⌕ Search

SEARCH · Search NASA

Results for “LDMS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

NERSC_Lightweight Distributed Metric Service (NERSC_LDMS) v4.4.2

Miscellany This LDMS Loftsman/Helm Chart horizontally scales LDMS daemons in order to achieve a 1Hz sample rate from over 5,000 nodes, collecting 38k metrics per minute on Perlmutter. This LMDS Configuration relies on already running `ldmsd` producers running on nodes, which produce metrics via sampler plugins. The Helm chart distributes the collection of metrics from producer acrross many aggregator and storage `ldmsd` daemons, ensuring no damon is overloaded and data loss is avoided.

Stile, John [Lawrence Berkeley National Laboratory↗

LDMS New Features for Deployment in Advanced Environments and Feedback for Operations

We describe how LDMS is being used to collect application data concurrent with system data and how the low-latency availability of this data for analysis can be used for real-time data analysis and feedback in order to support efficient, resilient, and reliable system operations. Finally, we will describe current related research areas.

Brandt, James Michael [Sandia National Laboratorie↗

LDMS Job Summary Pipeline

Explore the source record for details and available documents.

Schwaller, Benjamin [Sandia National Laboratories ↗

Integrated System and Application Continuous Performance Monitoring and Analysis Capability

Scientific applications run on high-performance computing (HPC) systems are critical for many national security missions within Sandia and the NNSA complex. However, these applications often face performance degradation and even failures that are challenging to diagnose. To provide unprecedented insight into these issues, the HPC Development, HPC Systems, Computational Science, and Plasma Theory & Simulation departments at Sandia crafted and completed their FY21 ASC Level 2 milestone entitled "Integrated System and Application Continuous Performance Monitoring and Analysis Capability." The milestone created a novel integrated HPC system and application monitoring and analysis capability by extending Sandia's Kokkos application portability framework, Lightweight Distributed Metric Service (LDMS) monitoring tool, and scalable storage, analysis, and visualization pipeline. The extensions to Kokkos and LDMS enable collection and storage of application data during run time, as it is generated, with negligible overhead. This data is combined with HPC system data within the extended analysis pipeline to present relevant visualizations of derived system and application metrics that can be viewed at run time or post run. This new capability was evaluated using several week-long, 290-node runs of Sandia's ElectroMagnetic Plasma In Realistic Environments ( EMPIRE ) modeling and design tool and resulted in 1TB of application data and 50TB of system data. EMPIRE developers remarked this capability was incredibly helpful for quickly assessing application health and performance alongside system state. In short, this milestone work built the foundation for expansive HPC system and application data collection, storage, analysis, visualization, and feedback framework that will increase total scientific output of Sandia's HPC users.

97 MATHEMATICS AND COMPUTING↗