Search NASA⌕ Search

SEARCH · Search NASA

Results for “workload”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Toward an IMU-Based Space Suit Motion Capture System

Spacesuits are complex engineering systems that sustain human health and enable performance outside Earth-like environments. These systems must support human mobility and physical workload demands while minimizing injury risk during extravehicular activity (EVA). Future EVA operations on the Lunar surface are expected to be more frequent and require higher physical workloads than previously during the ISS, Shuttle, and Apollo programs. To characterize the workloads and ergonomics needs a suit must support, the kinematics of the space suit must be measured during operationally-relevant tasks in ground analog environments. Kinematics capture of the suit is challenging for traditional optical motion capture (OMC) approaches due to marker occlusion, harsh lighting or environmental conditions, and tests with suit surrogates in outdoor field environments. To this end, engineers at NASA are developing the Augmented Suit Inverse Kinematics (ASIK) system, a complete motion capture method and inverse kinematics solver which relies solely on a network of wireless inertial measurement units (IMUs) attached to the major kinematic segments of the spacesuit. The ASIK modeling language allows for the simple inclusion of probabilistic priors such as suit size and shape or IMU poses. The ASIK system was tested in a 7-subject pilot study. Each subject donned NASA’s new prototype exploration spacesuit in the Active Response Gravity Offload System (ARGOS) facility at Johnson Space Center in Houston, TX. The suits were outfitted with 12 IMUs to estimate lower body and trunk kinematics. The suits were also outfitted with a set of reflective OMC markers, and traditional OMC data was collected and processed. Characterization of the ASIK-derived suit joint angles’ accuracy against an optical motion capture datum will be presented. Discussion of these results, as well as discussion of system calibration and nuances of mathematical observability, will be included.

IMU↗

Monitoring Airspace Complexity and Determining Contributing Factors

The national airspace has evolved over many years to accommodate increased traffic demand while simultaneously maintaining air travel as one of the safest forms of transportation. One of the reasons for this success is the ability of the air traffic control system and the operators to adapt and accommodate to situations that routinely disrupt normal operations. These situations may include: adverse weather, delays, early arrivals, equipment outages, and other factors that are outside the operators’ ability to control. These factors can lead to states where automation is unable to properly handle these issues, and therefore air traffic controllers and pilots have to intervene — ultimately increasing communication between operators resulting in higher workload. As controller workload increases to handle sub-optimal operating conditions, complexity increases. This is because, under these conditions humans are required to make tactical decisions in response to external factors. This results in a departure from the original strategic plan where operations would be more efficiently managed. Human operators manage airspace complexity under rigid regulations but in a constantly changing environment. The airspace is divided into sectors and the number of aircraft assigned to each controller is limited for safe handling. Some prior studies devised airspace complexity metrics in commercial aviation and related these metrics to controller workload. The upper bounds on the system load are pre-determined. Such bounds on complexity make for a safe system, but the system cannot scale and adapt to autonomous, dense, and heterogeneous traffic — including the many types of Unmanned Aerial Vehicles (UAVs) envisioned to be added to the operations. We hypothesize that, as traffic density and heterogeneity grow, and other key metrics change, there will be phase transitions at which the way traffic should be managed changes significantly. We offer a method for in-time detection of contributing factors that lead to phase transitions, characterized by increased complexity. To the best of our knowledge, there is no tool similar to ours that identifies such contributing factors or precursor patterns.

Precursor↗

Fatigue in Short-Haul Operations: Regulations and Research

Short-haul multi-segment flight operations are conducted within Part 117 table limits for Flight Time and Flight Duty Period, yet there are few data available regarding the impact on sleep, workload, and performance in these operations. The Civil Aerospace Medical Institute (CAMI) has invested in the planning and execution of a study that will result in data characterizing the effects of short-haul multi-segment flight operations on cumulative sleep loss and pilot workload across trip pairings. Scientific researchers from CAMI, in collaboration with NASA Ames Research Center, are currently running an operational field study to examine the human factors and elements of pilot performance associated with short-haul multi-segment flight operations to better understand the impact of sleep disruption and workload across trip pairings. The first half of this presentation will provide a brief overview of current duty and rest requirements and identify gaps in current operational research data. The second half of the presentation will focus on results from a focus group study which developed the scope of the larger field study, followed by an overview of the current field study.

short-haul↗

Lunar Terrain Vehicle (LTV) Remote Teleoperation Studies Under Four Lunar Communication Latencies

Remotely operating a lunar rover from Earth while subject to an Earth-Moon time delay of multiple seconds could result in a dangerous state where the roving vehicle is either damaged or lost, thereby potentially compromising an entire mission or series of missions. Providing the right capabilities to the remote operator to manage inherent communication latencies will be important for remote driving to be successful. NASA conducted two studies to investigate the average speed and number of kilometers per day that an operator on Earth could teleoperate a notional Artemis unpressurized rover with minimal remote operator capabilities under 0- and 4-second communication delays (April 2023 study) and 6- and 8-second delays (August 2023 study). A primary goal of these studies was to understand if an Artemis Lunar Terrain Vehicle (LTV) could cover 6 kilometers (km) in 24 hours when operated remotely. During the April 2023 evaluation, eight test operators used an in-house simulation of the lunar surface South Pole to teleoperate a NASA government reference LTV. Each operator received approximately 30 minutes of remote driving familiarization/training prior to their test run. Operators viewed the surrounding terrain via a single, rover mast-mounted, high-resolution camera with pan/tilt/zoom capabilities; continuous communication was provided throughout all testing. In the August 2023 evaluation, remote operators received approximately 3 hours of familiarization training in each latency, and the simulation environment provided remote operators with an operator-selected rate limiter to enable finer sensitivity in the hand controller and a predictive circle function to better assist operators with predicting the path the vehicle could take. All test operators were able to successfully navigate and drive through six different types of terrain and five planned traverse scenarios using natural lighting under all communication delays. Results for average speeds for each communication delay, computed by averaging the data from all test conditions for that latency and all operators, are shown in the table below. The average speed data was then used to derive the total time needed to cover 6 km, 8 km, and 20 km (distances relevant to LTV-SYS071 and -029 requirements). Remote operators drove slower and used the brake more frequently when subject to a communication latency as opposed to no communication latency. Subjective workload assessments revealed that while operating in a latency the overall workload significantly increased when compared to a 0-s delay with mental demand, frustration, and performance being the primary contributing factors. Driving strategies in the 0-s delay did not vary significantly among subjects; however, in the 4-s delay condition, three different driving strategies were identified. In the 6-s and 8-s latency conditions the operator’s use of the cruise control to maintain speed was more apparent. Additionally, over the course of the August study, the operator took advantage of the predictive circle indicator on the navigation display and over 95% of the operator’s navigation used the mast camera 180-degree panning function for ground truthing in terms of boulders and craters. Operators started to define more specific parameters in driving strategies for general operations. This consisted of setting the vehicle into a low-speed cruise mode of approximately 11.5 kph and noticing driving performance of the vehicle seemed to be much harder at slower speeds 0.4–0.8 kph; however, the vehicle was more responsive at speeds of 2.9–3.6 kph. Regardless of communication delay, operators used both the horizontal translation rails and the vehicle fenders as guides to predict a path for the vehicle through heavily concentrated terrain features. Test operators acknowledged that the teleoperations training for this study was substantially less than what an actual LTV remote operator will ultimately receive. They estimated a minimum of 20 to 100 hours spread across multiple days and weeks (e.g., strategies included immersion training over a 3-day period, to a short 8-week starter program) would be needed to get an operator ~ 60% proficient (i.e., able to complete a subset of remote driving tasks), to a yearlong program for full proficiency in remote driving tasks under all terrain types and natural lighting conditions. Remotely operating a vehicle on another planetary body while subject to communication latency is a complex task. Speed, distance covered, time spent driving, time spent navigating, brake usage and rock contacts are all affected by operator workload, driving strategies, workstation ergonomics and training. These studies provided a “first-look” answer to a potential system requirement (namely if a remote operator could cover a given distance in a given amount of time); however, considerable general knowledge was gained to begin to understand what it will take to make a successful lunar rover teleoperator.

LTV↗

Training the Powered-Lift Evaluation Pilot

This poster describes a project to prepare pilots for a study assessing novel aircraft automation concepts for electric Vertical Takeoff and Landing (eVTOL) aircraft using NASA’s Vertical Motion Simulator (VMS). By exploring the operational and learning challenges related to transitioning between forward flight and vertical landing, we seek to establish baselines of pilot workload and aircraft handling qualities across varying atmospheric conditions and automation states. The simulated eVTOL design differentiates flight control allocations as a function of airspeed across four speed ranges as the vehicle transitions between fully thrust-borne lift and wing-borne lift. As speed increases, side stick controls command: translational ground speeds, vertical and lateral acceleration, vertical rate, vertical flight path angle, and bank angle. This novel approach to flight control allocation creates a significant learning challenge for pilots. Since initial eVTOL aircraft may have limitations on hover capabilities, automation and flight guidance cues also vary with airspeed to provide efficient landing profiles while still providing cues suitable for cruise flight. The NASA team prepared the study pilots to follow these flight guidance cues along curved Required Navigation Performance (RNP) approaches and along 6o and 12o glide paths to energy-efficient assistive-hover landing and goarounds. The pre-VMS preparation sought to prepare pilots from diverse levels of experience and background. To do this, NASA researchers designed and developed a fixed-based, large field-ofview simulator with terrain, structures, and air traffic. With one day of combined classroom learning and skill development in the fixedbase simulator, pilots were largely able to fly the simulated eVTOL in the VMS with sufficient mastery to provide handling quality assessments using the Cooper-Harper Handling Qualities Rating and workload assessments through the Bedford Workload Scale.

AAM↗

Lunar Terrain Vehicle (LTV) Remote Teleoperation Studies Under Four Lunar Communication Latencies

Remotely operating a lunar rover from Earth while subject to an Earth-Moon time delay of multiple seconds could result in a dangerous state where the roving vehicle is either damaged or lost, thereby potentially compromising an entire mission or series of missions. Providing the right capabilities to the remote operator to manage inherent communication latencies will be important for remote driving to be successful. The National Aeronautics and Space Administration (NASA) conducted two studies to investigate the average speed and number of kilometers per day that an operator on Earth could teleoperate a notional Artemis unpressurized rover with minimal remote operator capabilities under 0- and 4-second communication delays (April 2023 study) and 6- and 8-second delays (August 2023 study). A primary goal of these studies was to understand if an Artemis Lunar Terrain Vehicle (LTV) could cover 6 kilometers (km) in 24 hours when operated remotely. During the April 2023 evaluation, eight test operators used an in-house simulation of the lunar surface South Pole to teleoperate a NASA government reference LTV. Each operator received approximately 30 minutes of remote driving familiarization/training prior to their test run. Operators viewed the surrounding terrain via a single, rover mast-mounted, high-resolution camera with pan/tilt/zoom capabilities; continuous communication was provided throughout all testing. In the August 2023 evaluation, remote operators received approximately 3 hours of familiarization training in each latency, and the simulation environment provided remote operators with an operator-selected rate limiter to enable finer sensitivity in the hand controller and a predictive circle function to better assist operators with predicting the path the vehicle could take. All test operators were able to successfully navigate and drive through six different types of terrain and five planned traverse scenarios using natural lighting under all communication delays. Results for average speeds for each communication delay, computed by averaging the data from all test conditions for that latency and all operators, are shown in the table below. The average speed data was then used to derive the total time needed to cover 6 km, 8 km, and 20 km. Remote operators drove slower and used the brake more frequently when subject to a communication latency as opposed to no communication latency. Subjective workload assessments revealed that while operating in a latency the overall workload significantly increased when compared to a 0-s delay with mental demand, frustration, and performance being the primary contributing factors. Driving strategies in the 0-s delay did not vary significantly among subjects; however, in the 4-s delay condition, three different driving strategies were identified. In the 6-s and 8-s latency conditions the operator’s use of the cruise control to maintain speed was more apparent. Additionally, over the course of the August study, the operator took advantage of the predictive circle indicator on the navigation display and over 95% of the operator’s navigation used the mast camera 180-degree panning function for ground truthing in terms of boulders and craters. Operators started to define more specific parameters in driving strategies for general operations. This consisted of setting the vehicle into a low-speed cruise mode of approximately 1–1.5 kph and noticing driving performance of the vehicle seemed to be much harder at slower speeds 0.4–0.8 kph; however, the vehicle was more responsive at speeds of 2.9–3.6 kph. Regardless of communication delay, operators used both the horizontal translation rails and the vehicle fenders as guides to predict a path for the vehicle through heavily concentrated terrain features. Test operators acknowledged that the teleoperations training for this study was substantially less than what an actual LTV remote operator will ultimately receive. They estimated a minimum of 20 to 100 hours spread across multiple days and weeks (e.g., strategies included immersion training over a 3-day period, to a short 8-week starter program) would be needed to get an operator ~ 60% proficient (i.e., able to complete a subset of remote driving tasks), to a yearlong program for full proficiency in remote driving tasks under all terrain types and natural lighting conditions. Remotely operating a vehicle on another planetary body while subject to communication latency is a complex task. Speed, distance covered, time spent driving, time spent navigating, brake usage and rock contacts are all affected by operator workload, driving strategies, workstation ergonomics and training. These studies provided a “firstlook” answer to a potential system requirement (namely if a remote operator could cover a given distance in a given amount of time); however, considerable general knowledge was gained to begin to understand what it will take to make a successful lunar rover teleoperator.

LTV↗

A Queuing Theory Approach to Pilot-Controller Coordination for m:N Operations

In recent years, attention and interest by industry and researchers has grown in a control paradigm for remotely piloted aircraft termed “m:N operations.” In an m:N operation, a team of m remote pilots in command (RIPCs) collaboratively manage the flights of N aircraft. A consequence of an m:N concept of operations is that the RPICs will have to switch attention from one aircraft to another and from one task to another. Previous research in m:N operations has focused on the workload experienced by an RPIC and their level of situation awareness on their flights. Researchers have found that RPIC workload and situation awareness are generally sensitive to increasing N, although NASA’s Multi-Vehicle (m:N) Working Group has suggested that the driver of workload/situation awareness is the number of exceptions requiring human intervention as opposed to the value of N itself. In any case, a natural antecedent of workload is task load. In this paper, queueing theory is applied to a 1:N Urban Air Mobility (UAM) air taxi operation in order to estimate pilot task load for managing radio communications with air traffic controllers (ATCs) under increasing N. An M/M/1 queueing system is used to model the RIPC’s servicing of calls and clearance requests (e.g., departure, arrival, or airspace transition) to ATC for the N aircraft. Important parameters for the queueing model are the task arrival rate and the average service time for task completion. Radio communication times from past human-in-the-loop simulation studies are used to measure service times for a 1:4 and 1:12 UAM operation and to interpolate service times for 4 < N < 12. A Monte Carlo method is then employed, using the measured and interpolated service times, to estimate arrival rate and related queueing statistics. The paper concludes by considering the estimated queuing statistics, particularly the RPIC’s utilization (i.e., proportion of time actively servicing tasks), the length of the task queue over time, and the implications for task-balanced system design.

task load↗

Communication Delays in Cislunar Space: A Lab Study Examining Human System Integration Architecture (HSIA) and Team Risk Concerns

BACKGROUND: Communication delays are an inherent challenge of space missions to the Moon and beyond. Past research studies showed that 50+ second delays adversely affected individual well-being, team cohesion, and overall task performance (Kintz et al., 2016; Larson et al., 2019). However, there is a dearth of research on the effects of shorter delays (on the order of 4-12 seconds one-way) that may present a more immediate challenge during the upcoming Artemis missions. Even these shorter delays could make it difficult or infeasible for ground control to provide real-time support to Artemis astronauts, especially in complex and time-critical tasks such as extra-vehicular activities (EVAs). Results from studies on longer or Mars-like delays cannot be directly applied to Artemis-like delays, due to the large differences in delay magnitude and task types between these mission categories. For example, real-time oversight and guidance from ground control is impossible under Mars-like delays but may be performed under Artemis-like delays, albeit with potentially high workload and communication difficulty. Thus, a better understanding of the effects of Artemis-like communication delays on collaborative task performance is needed. Additionally, there is a need to develop reliable task paradigms that can be used in future studies on communication delays. METHODS: This project will study the effects of Artemis-like communication delays on collaborative task performance in a simulated space-to-ground team task via a lunar Gateway-like interface prototype. Teams of two astronaut-like participants - one serving as “crewmember” and one as “flight controller” - will perform a spaceflight-relevant task under six delay conditions (i.e., 0, 4, 6, 8, 10, and 12 seconds). Audio, video, texting, and file transfer will be delayed between the participants to mimic lunar-like communication delays. Measures related to workload, situation awareness, system usability, task performance, team cohesion, well-being, and communication strategies will be collected. These measures will be compared across delay levels, participant roles, and off-nominal and nominal tasks. Participants will be given the choice to use any combination of video and text communication in order to study preferences in interaction modality. RESEARCH AIMS: One aim is to identify a delay level or range of levels at which there may be significant decrements in individual and team-based measures. Identifying such ranges may help design novel workload management or communication countermeasures for future space missions. Also, the spaceflight-relevant research tasks and corresponding communication delay technology developed as part of this effort may be used for future studies on communication delays. Results from this work are also expected to inform study scenarios in the Human Exploration Research Analog (HERA) Campaign.

S Upasani↗

ION Work Reduction Opportunity Realization Demonstration

The purpose of this research was to realize one of the advanced training work reduction opportunities first presented in the Idaho National Laboratory (INL) report, “Process for Significant Nuclear Work Function Innovation Based on Integrated Operations Concepts” (INL/EXT-21-64134) [1], with a nuclear power plant (NPP) research partner. Researchers modernized two trainings: (1) an accredited instructor-led training (ILT) overview course on Westinghouse DS 480-volt (V) circuit breakers to a multimedia-focused computer-based-training (CBT) learning module, and (2) an on-demand chaptered video on how to properly rack and un-rack a Westinghouse DS 480-V circuit breaker. These modernized work products were developed and implemented in a manner consistent with the industry guidelines found in Institution of Nuclear Power Operations (INPO) Teaching and Learning 23-001 [2]. Researchers calculated that the modernized accredited training course reduced the time necessary to prepare and deliver the training material by a factor of 8:1. The amount of time learners spend in class could be reduced by this same factor. In other words, if a course took 8 hours to deliver a class, the new CBT instruction would take just over 1 hour. The researchers noted that the requirement for any practicum training by the learners with the instructor(s) would remain in place. But through interviews with new and experienced learners, the researchers discovered that the confidence of these learners in performing the racking and un-racking of the circuit breaker improved as a result of using the new modernized CBT process. Additionally, the learners who tested the modernized work products enjoyed the modernized CBT and the learning video significantly more than current in-class learning methods. These are encouraging results for the nuclear industry, as this modernization of training can be applied to other classes and is scalable across the industry. In line with the Integrated Operations for Nuclear (ION) model, positive workload analysis supports the investment of resources in modernizing NPP training processes and infrastructure. Implementation of the advanced training technologies in this report is likely to result in substantive long-term workload benefits to instructors and learners and result in hard-dollar savings on contractor spends. Additionally, investment in these modernized training processes will result in improved learner proficiency. The results of this research can be applied to additional operator, technical, and general training topics to provide additional workload and learning benefits in addition to what was explored.

42 ENGINEERING↗

Evaluation of Best Practices in Mitigating Startup Costs on Leadership-Class Supercomputers

Supercomputers at Department of Energy (DOE) National Laboratories face a widening range of workloads, from traditional modeling and simulation to Artificial Intelligence model training or complex multi-stage workflows, and beyond. At DOE Leadership Computing Facilities like the Oak Ridge Leadership Computing Facility (OLCF), these workloads demand concurrent access to large portions of the supercomputer’s resources. Launching a job across massive supercomputers is challenging from the start; the file system struggles with a large backlog of metadata requests as tens of thousands of processes read thousands of the same files, and the compute job cannot start until this is completed. There are multiple existing approaches to calm this metadata storm, ranging from vendor-developed tools like sbcast to National Laboratory-developed tools like Spindle and Copper. In this paper, we benchmark and discuss three common approaches to improving compute job launch latencies on Frontier: Slurm’s sbcast tool, Spindle, and Copper. We evaluate these tools by measuring the launch latencies of four workloads: OSU Microbenchmark’s osu_init, Pynamic, Python import mpi4py, and Python import torch. We provide discussion of the results, highlighting data that meet expectations and that do not meet expectations.

Hagerty, Nick [ORNL] (ORCID:0000000330014414)↗

Evaluation of LLVM Flang for Production HPC Applications and Modern Fortran Features

In 2025, LLVM released its first Flang Fortran compiler version considered ready for widespread evaluation. We know of no published assessment of Flang compiling a workload- derived portfolio of high-performance computing (HPC) applications. We address this gap using workload data from the National Energy Research Scientific Computing Center (NERSC), which supports more than 10,000 scientists on approximately 1,000 projects. The NERSC workload analyses identify many Fortran components in heavily used applications. We selected 10 such packages with available source code. We compiled them with Flang 22.1.3 on NERSC’s Perlmutter system. Six compiled without code modifications, though some required build-system changes. Three compiled after minor source edits, mostly to address Fortran standard violations. One built only without OpenMP enabled. We evaluated seven additional packages selected for their use of, or enablement of, standard Fortran parallel features: multi-image execution and do concurrent. Six such codes compiled with most or all unit tests passing.

Rasmussen, Katherine↗

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq↗

Mixed-precision numerics in scientific applications: survey and perspectives

The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.

Graphics processing units↗

Operational Analytics Studies for ATLAS Distributed Computing: Data Popularity Forecast and Utilization of the WLCG Centers

Operational analytics is the direction of research related to the analysis of the current state of computing processes and the prediction of future states in order to anticipate imbalances and take timely measures to stabilize a complex system. There are two relevant areas in ATLAS Distributed Computing that are currently the focus of studies: user physics analysis including the forecast of popularity of data samples among users, and evaluating WLCG centers for their readiness to process user analysis payloads. Studying these areas is challenging due to the complexity involved, as it requires a comprehensive understanding of numerous boundary conditions typically found in large-scale distributed computing infrastructures. Forecasts of data popularity are problematic without the categorization of user tasks by their types (data transformation or physics analysis), which do not always appear on the surface but may induce noise, which introduces significant distortions for predictive analysis. Evaluating the WLCG resources by their analysis workloads is also a challenging task as it is necessary to find a balance between the workload of the resource, its performance, the waiting time for jobs on it, as well as the volume of jobs that it processes. This is especially difficult in a heterogeneous computing environment, where legacy resources are used along with modern high-performance machines. We will look at these areas of research in detail and discuss what tools and methods are used in our work, demonstrating results already obtained.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Supporting multiple hardware architectures at CMS: the integration and validation of POWER9

Computing resources in the Worldwide LHC Computing Grid (WLCG) have been based entirely on the x86 architecture for more than two decades. In the near future, however, heterogeneous non-x86 resources, such as ARM, POWER and Risc-V, will become a substantial fraction of the resources that will be provided to the LHC experiments, due to their presence in existing and planned world-class HPC installations. The CMS experiment, one of the four large detectors at the LHC, has started to prepare for this situation, with the CMS software stack (CMSSW) already compiled for multiple architectures. In order to allow for a production use, the tools for workload management and job distribution need to be extended to be able to exploit heterogeneous architectures. Profiting from the opportunity to exploit the first sizable IBM Power9 allocation available on Marconi100 HPC system at CINECA, CMS developed all the needed modifications to the CMS workload management system. After a successful proof of concept, a full physics validation has been performed in order to bring the system in production. The experiences are of very high value, when it comes to commissioning of the similar (even larger) Summit HPC system at Oak Ridge, where CMS is also expecting a resource allocation. Moreover the compute power of those systems is being provided also via GPUs and this represents an extremely valuable opportunity to exploit the offloading capability already implemented in CMSSW. The status of the current integration including the exploitation of the GPUs, the results of the validation as well as the future plans will be shown and discussed.

Boccali, Tommaso [INFN, Pisa]↗

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C↗

JACC.shared: Leveraging HPC Metaprogramming and Performance Portability for Computations That Use Shared Memory GPUs

In this work, we present JACC.shared, a new feature of Julia for ACCelerators (JACC), which is the performanceportable and metaprogramming model of the just-in-time and LLVM-based Julia language. This new feature allows JACC applications to leverage the high-performance computing (HPC) capabilities of high-bandwidth, on-chip GPU memory. Historically, exploiting high-bandwidth, shared-memory GPUs has not been a priority for high-level programming solutions. JACC.shared covers that gap for the first time, thereby providing a highlevel, portable, and easy-to-use solution for programmers to exploit this memory and supporting all current major accelerator architectures. Well-known HPC and AI workloads, such as multi/hyperspectral imaging and AI convolutions, have been used to evaluate JACC.shared on two exascale GPU architectures hosted by some of the most powerful US Department of Energy supercomputers: Perlmutter (NVIDIA A100) and Frontier (AMD MI250X). The performance evaluation reports speedup of up to 3.5× by adding only one line of code to the base codes, thus providing important accelerators in a simple, portable, and transparent way and elevating the programming productivity and performance-portability capabilities for Julia/JACC HPC, AI, and scientific applications.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗