Search NASA⌕ Search

SEARCH · Search NASA

Results for “workload”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Transitioning GlideinWMS, a multi domain distributed workload manager, from GSI proxies to tokens and other granular credentials

GlideinWMS is a distributed workload manager that has been used in production for many years to provision resources for experiments like CERN’s CMS, many Neutrino experiments, and the OSG. Its security model was based mainly on GSI (Grid Security Infrastructure), using X.509 certificate proxies and VOMS (Virtual Organization Membership Service) extensions. Even when other credentials, like SSH keys, were used to authenticate with resources, proxies were also added all the time, to establish the identity of the requestor and the associated memberships or privileges. This single credential was used for everything and was, often implicitly, forwarded wherever needed. The addition of identity and access tokens and the phase-out of GSI forced us to reconsider the security model of GlideinWMS, to handle multiple credentials which can differ in type, technology, and functionality. Both identity tokens and access tokens are supported. GSI proxies even if no more mandatory, are still used, together with various JWT (JSON Web Token) based tokens and other certificates. The functionality of the credentials, defined by issuer, audience, and scope, also differ: a credential can allow access to a computing resource, or can protect the GlideinWMS framework from tampering, or can grant read or write access to storage, can provide an identity for accounting or auditing, or can provide a combination of any the formers. Furthermore, the tools in use do not include automatic forwarding and renewal of the new credentials so credential lifetime and renewal requirements became part of the discussion as well. In this paper, we will present how GlideinWMS was able to change its design and code to respond to all these changes.

97 MATHEMATICS AND COMPUTING↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Using Apptainer in a Pilot-based Distributed Workload

GlideinWMS is a pilot and pressure-based workload manager for distributed scientific computing. Many experiments like CMS and Fermilab’s Neutrino experiments use it to provision elastic clusters for their analysis and simulations, split into close to a million concurrent jobs. Most user jobs require containers, and the pilots use Apptainer to set up the desired platform. For the pilots that run as regular batch jobs, Apptainer is safer, lighter, and easier to use than other containerization solutions. Many images used by the pilots are expanded SIF images distributed via the CernVM-FS: this combination is very efficient. At Fermilab, for example, we store on GitHub Dockerfiles that mimic the platform in the worker nodes of local clusters. GitHub workflows build and push the images to Docker Hub, and a service periodically pulls and converts them to the expanded SIF images in the CernVM-FS, so the scientists can find a familiar environment everywhere. Apptainer has also been used to run services inside the pilot jobs, like benchmarks that characterize the worker node being used, or a Triton Inference Server that allows sharing a GPU with all the jobs that run in parallel on a node.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

Time estimation as a secondary task to measure workload

Variation in the length of time productions and verbal estimates of duration was investigated to determine the influence of concurrent activity on operator time perception. The length of 10-, 20-, and 30-sec intervals produced while performing six different compensatory tracking tasks was significantly longer, 23% on the average, than those produced while performing no other task. Verbal estimates of session duration, taken at the end of each of 27 experimental sessions, reflected a parallel increase in subjective underestimation of the passage of time as the difficulty of the task performed increased. These data suggest that estimates of duration made while performing a manual control task provide stable and sensitive measures of the workload imposed by the primary task, with minimal interference.

Sandra G. Hart↗

A kinesthetic-tactual display concept for helicopter-pilot workload reduction

A kinesthetic-tactual (K-T) display concept is now under research and development (R & D) at the Ohio State University. It appears to offer considerable promise for useful application in helicopters by conveying control information via the sense of touch. This is a review of the overall R & D program including the original K-T display design, initial studies in automobile and fixed-wing vehicles, and feasibility experiments in a helicopter simulator. In addition to investigations of control and potential workload reduction, present efforts are directed toward establishing optimal design requirements for K-T helicopter displays. Potential applications, modes of usage, and the kinds of information that may be displayed in helicopter applications are discussed along with a brief forecast of future R & D. A brief description of the latest multi-axis laboratory prototype K-T display is also provided.

Gilson, R. D.↗

The effects of participatory mode and task workload on the detection of dynamic system failures

The ability of operators to detect step changes in the dynamics of control systems is investigated as a joint function of, (1) participatory mode: whether subjects are actively controlling those dynamics or are monitoring an autopilot controlling them, and (2) concurrent task workload. A theoretical analysis of detection in the two modes identifies factors that will favor detection in either mode. Three subjects detected system failures in either an autopilot or manual controlling mode, under single-task conditions and concurrently with a subcritical tracking task. Latency and accuracy of detection were assessed and related through a speed accuracy tradeoff representation. It was concluded that failure detection performance was better during manual control than during autopilot control, and that the extent of this superiority was enhanced as dual-task load increased. Ensemble averaging and multiple regression techniques were then employed to investigate the cues utilized by the subjects in making their detection decisions.

Wickens, C. D.↗

Secondary visual workload capability with primary visual and kinesthetic-tactual displays

Subjects performed a cross-adaptive tracking task with a visual secondary display and either a visual or a quickened kinesthetic-tactual (K-T) primary display. The quickened K-T display resulted in superior secondary task performance. Comparisons of secondary workload capability with integrated and separated visual displays indicated that the superiority of the quickened K-T display was not simply due to the elimination of visual scanning. When subjects did not have to perform a secondary task, there was no significant difference between visual and quickened K-T displays in performing a critical tracking task.

Gilson, R. D.↗

NASA TLA workload analysis support. Volume 2: Metering and spacing studies validation data

Four sets of graphic reports--one for each of the metering and spacing scenarios--are presented. The complete data file from which the reports were generated is also given. The data was used to validate the detail task of both the pilot and copilot for four metering and spacing scenarios. The output presents two measures of demand workload and a report showing task length and task interaction.

Sundstrom, J. L.↗

Visual scanning behavior and mental workload in aircraft pilots

This paper describes an experimental paradigm and a set of preliminary results which demonstrate a relationship between the level of performance on a skilled man-machine control task, the skill of the operator, the level of mental difficulty induced by an additional task imposed on the basic control task, and visual scanning performance. During a constant, simulated piloting task, visual scanning of instruments was found to vary as a function of the level of difficulty of a verbal loading task. The average dwell time of each fixation on the pilot's primary instrument increased as a function of the loading. The scanning behavior was also a function of the estimated skill level of the pilots, with novices being affected by the loading task much more than experts. The results suggest that visual scanning of instruments in a controlled task may be an indicator of both workload and skill.

Tole, J. R.↗

Comparative evaluation of twenty pilot workload assessment measure using a psychomotor task in a moving base aircraft simulator

A comparison of the sensitivity and intrusion of twenty pilot workload assessment techniques was conducted using a psychomotor loading task in a three degree of freedom moving base aircraft simulator. The twenty techniques included opinion measures, spare mental capacity measures, physiological measures, eye behavior measures, and primary task performance measures. The primary task was an instrument landing system (ILS) approach and landing. All measures were recorded between the outer marker and the middle marker on the approach. Three levels (low, medium, and high) of psychomotor load were obtained by the combined manipulation of windgust disturbance level and simulated aircraft pitch stability. Six instrument rated pilots participated in four seasons lasting approximately three hours each.

Connor, S. A.↗